arrow
返回

Quality Anomaly Detection Using Predictive Techniques: An Extensive Big Data Quality Framework for Reliable Data Analysis

delete2023-01-01
delete5
delete
OA
AI
W
Widad Elouataoui *
S
Saida El Mendili
Y
Youssef Gahi
DOI:10.1109/ACCESS.2023.3317354delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The increasing reliance on Big Data analytics has highlighted the critical role of data quality in ensuring accurate and reliable results. Consequently, organizations aiming to leverage the power of Big Data recognize the crucial role of data quality as an integral component. One notable type of data quality anomaly observed in big datasets is the presence of outlier values. Detecting and addressing these outliers have become a subject of interest across diverse domains, leading to the development of numerous anomaly detection approaches. Although anomaly detection has witnessed a proliferation of practices in recent years, a significant gap remains in addressing anomalies related to the other aspects of data quality. Indeed, while most approaches focus on identifying anomalies that deviate from the expected patterns, they do not consider irregularities in data quality, such as missing, incorrect, or inconsistent data. Moreover, most of approaches are domain-correlated and lack the capability to detect anomalies in a generic manner. Thus, we aim through this paper to address this gap in the field and provide a holistic and effective solution for Big Data quality anomaly detection. To achieve this, we suggest a novel approach that allows a comprehensive detection of Big Data quality anomalies related to six quality dimensions: Accuracy, Consistency, Completeness, Conformity, Uniqueness, and Readability. Moreover, the framework allows for sophisticated detection of generic data quality anomalies through the implementation of an intelligent anomaly detection model without any correlation to a specific field. Furthermore, we introduce and measure a new metric called Quality Anomaly Score, which refers to the degree of anomalousness of the quality anomalies of each quality dimension and the entire dataset. Through the implementation and evaluation of our framework, the suggested framework has achieved an accuracy score of up to 99.91% and an F1-score of 98.07%.
Keyword:
Data integrity
Big Data
Anomaly detection
Organizations
Reliability
Measurement
Data models
big data
big data quality
data quality dimensions
quality anomaly score

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

I
ibn tofail university of kenitra
学者数:
2.5K
论文数: 1.6K
被引数: 1
引用论文

引用论文

Extended Isolation Forest扩展隔离林
err2021-04-01
err190
errOAAI
errHariri, Sahand; Kind, Matias Carrasco; Brunner, Robert J.
err分享
err收藏
Fast outlier detection for very large log data
err2011-08-01
err31
PREAI
errKim, Seung; Cho, Nam Wook; Kang, Bokyoung; Kang, Suk-Ho
err分享
err收藏
err分享
err收藏
An Adaptable Big Data Value Chain Framework for End-to-End Big Data Monetization
err2020-11-23
err27
errOAAI
errFaroukhi, Abou Zakaria; El Alaoui, Imane; Gahi, Youssef; Amine, Aouatif
err分享
err收藏
History of Rabies
err1975-01-01
err0
PREAI
errJames H. Steele
err分享
err收藏
err分享
err收藏
Rare-earth elements and yttrium distributions in mangrove coastal water systems: The western Gulf of Thailand
err2005-08-01
err0
PREAI
errP. Censi; S. E. Spoto; G. Nardone; F. Saiano; R. Punturo; S. I. Di Geronimo; S. Mazzola; A. Bonanno; B. Patti; M. Sprovieri; D. Ottonello
err分享
err收藏
学者 查看更多内容