返回
Fuzzy C-Means-based Isolation Forest
DOI:10.1016/j.asoc.2021.107354.png)
摘要
En 中文
The problem of finding anomalies (outliers) in databases is one of the most important issues in modern data analysis. One of the reasons is the occurrence of this issue in almost every type of database, including numerical, categorical, time, mixed, or graphic data. There are currently many methods often dedicated to specific data analysis. Finally, this topic is extremely interesting per se, as a research problem that intrigues researchers. One of the classic methods of data analysis dedicated to finding the anomalies in the data is Isolation Forest. However, this method, with a few exceptions, has not been modified from the time of its first publication, and, in particular, it has not yet appeared in combination with the typical fuzzy methods used for grouping such as Fuzzy C-Means (FCM) clustering. In this study, we thoroughly analyze this approach, as well as several related ones. We examine the possibilities of this technique and analyze it in detail for characteristics of data (database size, number of attributes, records, their type, etc.). It is worth noting that FCM allows to obtain membership grades of elements forming Isolation Forest nodes to clusters on the basis of which these nodes are built. Hence, at the stage of calculating the anomaly scores, this information is effectively used, in particular to express how much a given element may belong to a group of similar elements, which can be inferred from the characteristics of the cluster in which it lies. In this study, we propose a set of methods enhancing the Isolation Forest on a basis of Fuzzy C-Means. The results of numerical experiments carried using 27 various datasets and reported in this paper lead us to the conclusion that FCM can play a pivotal role in an enhancement of Isolation Forest approach and raises up the values of particular measures of effectiveness of the anomaly detection methods. (C) 2021 Elsevier B.V. All rights reserved.
Keyword:
Fuzzy C-Means
Isolation Forest
Anomaly detection
Outliers detection
Fraud detection
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.6
论文数:
1.4W
被引数:
4.8W
机构
引用论文
Base‐catalyzed polymerization of maleimide and some derivatives and related unsaturated carbonamides马来酰亚胺和一些衍生物及相关不饱和碳酰胺的碱催化聚合


