arrow
Return

Document length normalization

delete1996-09-01
delete97
delete
OA
AI
A
Amit Singhal *
G
Gerard Salton
M
Mandar Mitra
C
Chris Buckley
DOI:10.1016/0306-4573(96)00008-8delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In the TREC collection-a large full-text experimental text collection with widely varying document lengths-we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique-pivoted cosine normalization-that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness. Copyright (C) 1996 Elsevier Science Ltd
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

No organization information available