返回
Fraud detection in the distributed graph database
DOI:10.1007/s10586-022-03540-3.png)
摘要
En 中文
Over the last few decades, graphs have become increasingly important in many applications and domains for managing Big data. Big data analysis in a graph database is described as an analysis of exponentially increasing massive interconnected data concerning time. However, analyzing big connected data in social networks and synthetic identity detection is challenging. In previous approaches, fraud detection has been done on the complete graph data, which is a time-consuming process and will create bottlenecks while query execution. To overcome the issue, this paper proposes a new fraud detection technique to unveil synthetic identities involved in the Panama Paper leak dataset (unprecedented leak of 11.5 m data from the database of the world's fourth-biggest offshore law arm, Mossack Fonseca) using a Node rank-based fraud detection algorithm by integrating distributed data profiling techniques on a minimized graph by minimizing the least influential nodes. The proposed model is verified on the three nodes cluster to improve data scalability, reduce the query execution time by an average of 30-36% and finally reduce the fraud detection time by 18.2%.
Keyword:
Neo4j
Graph database
Distributed system
Community fraud detection
Node rank
Minimized graph
Data profiling
NRFD
期刊
C
IF:
4.1
论文数:
5.0K
被引数:
7.5K

