返回
Data Mining Scheme for Globally Distributed Big Data
DOI:10.1080/08874417.2025.2564437.png)
摘要
En 中文
Globalization has led to a massive increase in the amount and variety of data generated worldwide as businesses strive to cater to the international market. However, current data mining techniques struggle to analyze data that is rapidly generated and stored at multiple remote locations. This is because data typically needs to be centrally aggregated for analysis, and transferring data worldwide increases processing costs and time. This paper proposes a scheme that includes a system architecture, communication flows, and a simplification method to improve the performance of machine learning techniques by simplifying globally distributed data without centralizing it. The proposed method reduces the size of datasets while improving the performance of the analysis. The method can also be used for making real-time decisions based on distributed big data in a rapidly changing business environment. By analyzing Uber pickup and Airbnb datasets using K-means clustering, this research shows that the results of our simplified datasets are closer to those of the original datasets than those of the traditional sampling datasets. Additionally, the size of the simplified datasets is smaller than that of the sampling datasets. Overall, this method offers a significant advantage by requiring less data for analysis compared to existing sampling and distributed computing techniques, with demonstrated computational efficiency and theoretical guarantees.
Keyword:
K-means
clustering analysis
distributed big data
data mining
data simplification
期刊
IF:
4.2
论文数:
205
被引数:
3.1K
机构
引用论文
暂无论文信息

