返回
Two-Stage Document Length Normalization for Information Retrieval
DOI:10.1145/2699669.png)
摘要
En 中文
The standard approach for term frequency normalization is based only on the document length. However, it does not distinguish the verbosity from the scope, these being the two main factors determining the document length. Because the verbosity and scope have largely different effects on the increase in term frequency, the standard approach can easily suffer from insufficient or excessive penalization depending on the specific type of long document. To overcome these problems, this article proposes two-stage normalization by performing verbosity and scope normalization separately, and by employing different penalization functions. In verbosity normalization, each document is prenormalized by dividing the term frequency by the verbosity of the document. In scope normalization, an existing retrieval model is applied in a straightforward manner to the prenormalized document, finally leading us to formulate our proposed verbosity normalized (VN) retrieval model. Experimental results carried out on standard TREC collections demonstrate that the VN model leads to marginal but statistically significant improvements over standard retrieval models.
Keyword:
Algorithms
Experimentation
Performance
Theory
Verbosity normalization
scope normalization
document length normalization
retrieval heuristics
term frequency
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.1
论文数:
1.2K
被引数:
4.7K
机构
暂无机构信息
引用论文
INITIATIVE TO IMPROVE KNOWLEDGE OF NUTRITION IN PEDIATRIC FAMILIES IN RENAL CLINIC改善肾科诊所儿科家庭营养知识水平的倡议
Putting perinatal attention-deficit/hyperactivity disorder in context: intersecting across identities, diagnoses, and generations将围产期注意缺陷/多动障碍置于背景中:跨越身份、诊断和代际的交叉

