arrow
Return

Improving Data Cleaning by Learning From Unstructured Textual Data

delete2025-01-01
delete0
delete
OA
AI
R
Rihem Nasfi *
G
Guy De Tré
A
Antoon Bronselaer
DOI:10.1109/ACCESS.2025.3543953delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In data analysis, a significant amount of erroneous or incomplete data can hinder informed organizational decisions prompting the need for automated data cleaning. Leveraging successful artificial intelligence techniques across various domains, several initiatives have introduced machine learning models to tackle these data-quality related issues. In this paper, we study a novel strategy to enhance data cleaning using a machine learning-based model trained on unstructured data. The study involves training a machine learning model with instances that satisfy the dataset constraints (such as functional dependencies), and examining its impact on efficiency in repairing dirty data. We also suggest improving data repair by comparing the probabilities of machine learning model predictions against a threshold, replacing the actual data with higher certainty predictions. Moreover, a hybrid data cleaning approach can enhance the effectiveness of our proposed method, resulting in higher repair quality. Experiments demonstrate promising results in terms of precision (0.88) and recall (0.94), particularly when the training data are clean from constraint violations and sufficiently large.
Keywords:
Cleaning
Data models
Maintenance engineering
Predictive models
Accuracy
Numerical models
Machine learning
Transformers
Training
Soft sensors
Data cleaning
data quality
data constraints
machine learning
natural language processing

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

G
Ghent University
Scholars:
5.2W
Papers: 4.5W
Citations: 5.5W