Return
Document length normalization
DOI:10.1016/0306-4573(96)00008-8.png)
Abstract
En 中文
In the TREC collection-a large full-text experimental text collection with widely varying document lengths-we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique-pivoted cosine normalization-that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness. Copyright (C) 1996 Elsevier Science Ltd
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
Organization
No organization information available

