Return
Anomaly detection in data warehouse using machine learning techniques
DOI:10.1016/j.cie.2026.112365.png)
Abstract
En 中文
The exponential growth of data utilization has made data warehouses a crucial component for storing vast amounts of information. However, data stored in these repositories may contain anomalies, which can adversely impact the reliability and quality of analytical results. Detecting anomalies in data warehouses has thus become a critical concern to ensure data integrity. In this paper, we propose a new unsupervised approach that describes a warehouse table with a family of elementary detectors and makes them work together through a common severity scale. Deviations along the principal directions of the numerical attributes, marginal deviations, rarity of the signs, unexpected combinations of categorical modalities and conditional deviations of a numerical attribute inside a category are each converted into a severity by means of an empirical fence estimated on a reference set that the method cleans by itself. The score of a record is the largest severity observed among all the detectors, and a single global threshold is calibrated on this aggregated score, which takes into account the multiplicity of the detectors that are evaluated simultaneously. No label is used at any stage of the calibration, and every alert is delivered with the name of the detector that produced it. We demonstrate the effectiveness of our approach on a real dataset extracted from an e-commerce company’s data warehouse. The approach reaches a precision of 93.98%, a recall of 95.75%, an F-measure of 94.85% and a false positive rate of 0.125%, and it processes the whole test set in just under two minutes on an ordinary laptop without any graphics accelerator. These results make our approach a reliable and inexpensive tool for accurate anomaly detection in data warehouses across diverse domains and applications
Journal
C
IF:
6.5
Papers:
499
Citations:
0
Organization
Cited Papers
No cited papers available

