arrow
Return

A distributed incremental in; mation acquisition model; large-scale text data

delete2017-12-21
delete3
PRE
AI
孙胜涛 (Shengtao Sun)
宫继兵 (Jibing Gong) *
A
Albert Y. Zomaya
A
Aizhi Wu
DOI:10.1007/s10586-017-1498-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Timely discovering and acquiring in; mation from incremental data on the Internet is a hot topic in a big data era. This paper presents a distributed incremental in; mation acquisition model; large-scale text data. To obtain a lower false positive rate and higher efficiency of the traditional Bloom filter, a distributed multidimensional Bloom filter is designed and proposed to cope with the deduplication of large-scale Web URL text data. Three methods related to Bloom filter were compared based on the false positive rate and response efficiency. The results show that the distributed incremental in; mation acquisition model; large-scale text data can achieve a high duplicate removal rate with a lower false positive rate.
Keywords:
Big data analytics
Deduplication of large-scale text data
Distributed incremental in
mation acquisition model
Distributed multidimensional bloom filter
False positive rate
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
Papers:
5.0K
Citations:
7.5K

Organization

Y
Yanshan University
Scholars:
1.7W
Papers: 1.1W
Citations: 1.3W
U
University of Sydney
Scholars:
6.5W
Papers: 6.2W
Citations: 90