Return
Improving Data Cleaning by Learning From Unstructured Textual Data
DOI:10.1109/ACCESS.2025.3543953.png)
Abstract
En 中文
In data analysis, a significant amount of erroneous or incomplete data can hinder informed organizational decisions prompting the need for automated data cleaning. Leveraging successful artificial intelligence techniques across various domains, several initiatives have introduced machine learning models to tackle these data-quality related issues. In this paper, we study a novel strategy to enhance data cleaning using a machine learning-based model trained on unstructured data. The study involves training a machine learning model with instances that satisfy the dataset constraints (such as functional dependencies), and examining its impact on efficiency in repairing dirty data. We also suggest improving data repair by comparing the probabilities of machine learning model predictions against a threshold, replacing the actual data with higher certainty predictions. Moreover, a hybrid data cleaning approach can enhance the effectiveness of our proposed method, resulting in higher repair quality. Experiments demonstrate promising results in terms of precision (0.88) and recall (0.94), particularly when the training data are clean from constraint violations and sufficiently large.
Keywords:
Cleaning
Data models
Maintenance engineering
Predictive models
Accuracy
Numerical models
Machine learning
Transformers
Training
Soft sensors
Data cleaning
data quality
data constraints
machine learning
natural language processing

