arrow
Return

PAACDA: Comprehensive Data Corruption Detection Algorithm

delete2023-01-01
delete0
delete
OA
AI
C
Charvi Bannur
C
Chaitra Bhat
K
Kushagra Singh
S
Shrirang Ambaji Kulkarni *
M
Mrityunjay Doddamani
DOI:10.1109/ACCESS.2023.3253022delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the advent of technology, data and its analysis are no longer just values and attributes strewn across spreadsheets, they are now seen as a stepping stone to bring about revolution in any significant field. Data corruption can be brought about by a variety of unethical and illegal sources, making it crucial to develop a method that is highly effective to identify and appropriately highlight the various corrupted data existing in the dataset. Detection of corrupted data, as well as recovering data from a corrupted dataset, is a challenging problem. This requires utmost importance and if not addressed at earlier stages may pose problems in later stages of data processing with machine or deep learning algorithms. In the following work we begin by introducing the PAACDA: Proximity based Adamic Adar Corruption Detection Algorithm and consolidating the results whilst particularly accentuating the detection of corrupted data rather than outliers. Current state of the art models, such as Isolation forest, DBSCAN also called Density-Based Spatial Clustering of Applications with Noise and others, are reliant on fine-tuning parameters to provide high accuracy and recall, but they also have a significant level of uncertainty when factoring the corrupted data. In the present work, the authors look into the most niche performance issues of several unsupervised learning algorithms for linear and clustered corrupted datasets. Also, a novel PAACDA algorithm is proposed which outperforms other unsupervised learning benchmarks on 15 popular baselines including K-means clustering, Isolation forest and LOF (Local Outlier Factor) with an accuracy of 96.35% for clustered data and 99.04% for linear data. This article also conducts a thorough exploration of the relevant literature from the previously stated perspectives. In this research work, we pinpoint all the shortcomings of the present techniques and draw direction for future work in this field.
Keywords:
Anomaly detection
Clustering algorithms
Data integrity
Data models
Computer science
Biological system modeling
Unsupervised learning
Adamic Adar algorithm
corrupted datasets
outlier detection
probabilistic models
statistical models
unsupervised learning

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

N
national institute of engineering (nie)
Scholars:
171
Papers: 158
Citations: 0
I
indian institute of technology (iit) - mandi
Scholars:
893
Papers: 761
Citations: 5
I
indian institute of technology system (iit system)
Scholars:
9.5W
Papers: 9.9W
Citations: 93
researcher View more organizations
Cited Papers

Cited Papers

A Review of Local Outlier Factor Algorithms for Outlier Detection in Big Data Streams
err2020-12-29
err161
errOAAI
errAlghushairy, Omar; Alsini, Raed; Soule, Terence; Ma, Xiaogang
errShare
errSave
A Flying Squirrel Search Optimization for MPPT Under Partial Shaded Photovoltaic System
err2021-08-01
err0
PREAI
errNagendra Singh; Krishna Kumar Gupta; Sanjay K. Jain; Niraj Kumar Dewangan; Pallavee Bhatnagar
errShare
errSave
ECOD: Unsupervised Outlier Detection Using Empirical Cumulative Distribution Functions
err2023-12-01
err128
errOAAI
errLi, Zheng; Zhao, Yue; Hu, Xiyang; Botta, Nicola; Ionescu, Cezar; Chen, George H.
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
A Novel Outlier Detection Method for Multivariate Data
err2022-09-01
err33
errOAAI
errAlmardeny, Yahya; Boujnah, Noureddine; Cleary, Frances
errShare
errSave
errShare
errSave
researcher View more