Log In Sign Up

Contextualization of Big Data Quality: A framework for comparison

by   Mostafa Mirzaie, et al.

With the advent of big data applications and the increasing amount of data being produced in these applications, the importance of efficient methods for big data analysis has become highly evident. However, the success of any such method will be hindered should the data lacks the required quality. Big data quality assessment is therefore a major requirement for any organization or business that use big data analytics for its decision making. On the other hand, using contextual information is advantageous in many analysis tasks in various domains, e.g. user behavior analysis in the social networks. However, the big data quality assessment has benefited less from this potential. There is a vast variety of data sources in the big data domain that can be utilized to improve the quality evaluation of big data. Including contextual information provided by these sources into the big data quality assessment process is an emerging trend towards more advanced techniques aimed at enhancing the performance and accuracy of quality assessment. This paper presents a context classification framework for big data quality, categorizing the context features into four primary dimensions: 1) context category, 2) data source type that contextual features come from, 3) discovery and extraction method of context, and 4) the quality factors affected by the contextual data. The proposed model introduces new context features and dimensions that need to be taken into consideration in quality assessment of big data. The initial evaluation demonstrates that the model is more understandable, more comprehensive, richer, and more useful compared to existing models.


State of the Art on the Quality of Big Data: A Systematic Literature Review and Classification Framework

One of the most significant problems of Big Data is to extract knowledge...

Characterizing Big Data Management

Big data management is a reality for an increasing number of organizatio...

Aware Adoption of AI: from Potential to Reusable Value

Artificial Intelligence (AI) provides practical advantages in different ...

Ontology-based Design of Experiments on Big Data Solutions

Big data solutions are designed to cope with data of huge Volume and wid...

Ontological Multidimensional Data Models and Contextual Data Qality

Data quality assessment and data cleaning are context-dependent activiti...

Big Data Fusion to Estimate Fuel Consumption: A Case Study of Riyadh

Falling oil revenues and rapid urbanization are putting a strain on the ...

Survey of Big Data sizes in 2021

The modern increase in data production is driven by multiple factors, an...