arrow
Return

Collection statistics for fast duplicate document detection

delete2002-04-01
delete162
PRE
AI
A
Abdur Chowdhury *
O
Ophir Frieder
M
M. Catherine McCabe
DOI:10.1145/506309.506311delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present a new algorithm for duplicate document detection that uses collection statistics. We compare our approach with the state-of-the-art approach using multiple collections. These collections include a 30 MB 18,577 web document collection developed by Excite@ Home and three NIST collections. The first NIST collection consists of 100 MB 18,232 LA-Times documents, which is roughly similar in the number of documents to the Excite@ Home collection. The other two collections are both 2 GB and are the 247,491-web document collection and the TREC disks 4 and 5-528,023 document collection. We show that our approach called I-Match, scales in terms of the number of documents and works well for documents of all sizes. We compared our solution to the state of the art and found that in addition to improved accuracy of detection, our approach executed in roughly one-fifth the time.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Information Systems cover
ACM Transactions on Information Systems
IF:
9.1
Papers:
1.2K
Citations:
4.7K

Organization

No organization information available