Return
Correlation-based feature partition regression method for unsupervised anomaly detection
DOI:10.1007/s10489-022-03247-3.png)
Abstract
En 中文
Anomaly detection problem has been extensively studied in a variety of application domains, where the data tags are difficult to obtain. Most unsupervised algorithms rely on some notions such as distance and density to detect anomalies. However, the performance of such algorithms is easier to decrease as the dimension of the datasets increases. Some studies which use features as pseudo-labels for prediction detect anomalies according to the deviation value of the prediction model. Even so, the improvement of model performance is still restricted to ignoring the correlation between feature attributes. In this paper, we propose a correlation-based feature partition regression prediction method called CFPR, which can alleviate the adverse effects of dataset dimensions and irrelevant attributes on model performance to a certain extent. According to the correlation between the features, the high-dimensional datasets will be divided into multiple feature subspaces. In each subspace, the feature with the highest correlation coefficient will be conducted as a pseudo-label. After that, we use the remaining features as the prediction attributes to train a supervised regression prediction model. We can calculate the anomaly score of each sample in the subspace according to the difference between the regression prediction value and the true value of the pseudo-label. Furthermore, we define a weighting strategy based on the level of correlation in the subspace integration stage to obtain the final anomaly score ranking table. Extensive experiments on twenty-eight UCI public datasets show that the CFPR performs better than several state-of-art anomaly algorithms at the AUC metric.
Keywords:
Unsupervised anomaly detection
Pseudo labeling
Feature correlation partition
Weighted regression predict
Journal
IF:
3.5
Papers:
7.5K
Citations:
1.7W

