arrow
Return

Removing DUST Using Multiple Alignment of Sequences

delete2015-08-01
delete7
delete
OA
AI
M
Marco Cristo
E
Edleno Silva de Moura
D
da Silva, Altigran
DOI:10.1109/TKDE.2015.2407354delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
A large number of URLs collected by web crawlers correspond to pages with duplicate or near-duplicate contents. To crawl, store, and use such duplicated data implies a waste of resources, the building of low quality rankings, and poor user experiences. To deal with this problem, several studies have been proposed to detect and remove duplicate documents without fetching their contents. To accomplish this, the proposed methods learn normalization rules to transform all duplicate URLs into the same canonical form. A challenging aspect of this strategy is deriving a set of general and precise rules. In this work, we present DUSTER, a new approach to derive quality rules that take advantage of amulti-sequence alignment strategy. We demonstrate that a fullmulti-sequence alignment of URLs with duplicated content, before the generation of the rules, can lead to the deployment of very effective rules. By evaluating our method, we observed it achieved larger reductions in the number of duplicate URLs than our best baseline, with gains of 82 and 140.74 percent in two different web collections.
Keywords:
Web technology
web crawling and normalization rules
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.8K
Citations:
3.2W

Organization

U
universidade federal de amazonas
Scholars:
2.8K
Papers: 1.6K
Citations: 1