Amina Onyeabor, Grace and Ta'a, Azman (2018) Data quality evaluation framework for big data. i-manager’s Journal on Cloud Computing, 5 (2). p. 27. ISSN 2349-6835
PDF (front page)
Restricted to Registered users only Download (56kB) | Request a copy |
Abstract
Data is an important asset in all business organizations of today. Thus the results of its poor quality can be very grievous leading to erroneous insights. Therefore, Data Quality (DQ) needs to be evaluated before the analysis of any Big Data (BD). The evaluation of DQ in BD is challenging. Given the enormous datasets that are of varied format fashioned at a rapid speed, it is impossible to use the traditional methods of evaluating DQ in BD. Rather, there is a requirement of strategies and devices for the assessment and evaluation of DQ in BD in a rapid and more efficient manner. However, assessing the quality of data on the whole BD can be very expensive. In addition, there is also a need for improvement in data transformation activities of BD. This paper proposes a framework for DQ evaluation with the application of data sampling technique on BD sets from different data sources reducing the size of the data to samples representing the population of the BD sets. The Bag of Little Bootstrap (BLB) sampling technique will be used. The target Data Quality Dimensions (DQDs) to be used in this paper are completeness, consistency, and accuracy. In addition, the DQDs will be measured using different metric functions relevant to the DQDs. This will be done before and after an improved data transformation techniques to check the improvement of DQ in BD.
Item Type: | Article |
---|---|
Uncontrolled Keywords: | Big Data, Data Sampling, Data Transformation, Data Quality Evaluation. |
Subjects: | Q Science > QA Mathematics > QA75 Electronic computers. Computer science |
Divisions: | School of Computing |
Depositing User: | Mrs. Norazmilah Yaakub |
Date Deposited: | 09 Sep 2020 03:03 |
Last Modified: | 09 Sep 2020 03:03 |
URI: | https://repo.uum.edu.my/id/eprint/27441 |
Actions (login required)
View Item |