arrow
返回

Set-based vector model:: An efficient approach for correlation-based ranking

delete2005-10-01
delete29
PRE
AI
B
Bruno Pôssas
N
Nívio Ziviani
W
Wagner Meira
B
Berthier Ribeiro‐Neto
DOI:10.1145/1095872.1095874delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This work presents a new approach for ranking documents in the vector space model. The novelty lies in two fronts. First, patterns of term co-occurrence are taken into account and are processed efficiently. Second, term weights are generated using a data mining technique called association rules. This leads to a new ranking mechanism called the set-based vector model. The components of our model are no longer index terms but index termsets, where a termset is a set of index terms. Termsets capture the intuition that semantically related terms appear close to each other in a document. They can be efficiently obtained by limiting the computation to small passages of text. Once termsets have been computed, the ranking is calculated as a function of the termset frequency in the document and its scarcity in the document collection. Experimental results show that the set-based vector model improves average precision for all collections and query types evaluated, while keeping computational costs small. For the 2-gigabyte TREC-8 collection, the set-based vector model leads to a gain in average precision figures of 14.7% and 16.4% for disjunctive and conjunctive queries, respectively, with respect to the standard vector space model. These gains increase to 24.9% and 30.0%, respectively, when proximity information is taken into account. Query processing times are larger but, on average, still comparable to those obtained with the standard vector model (increases in processing time varied from 30% to 300%). Our results suggest that the set-based vector model provides a correlation-based ranking formula that is effective with general collections and computationally practical.
Keyword:
information retrieval models
association rule mining
weighting index term co-occurrences
data mining
correlation-based ranking

期刊

ACM Transactions on Information Systems 封面图
ACM Transactions on Information Systems
IF:
9.1
论文数:
1.2K
被引数:
4.7K

机构

暂无机构信息
引用论文

引用论文

err分享
err收藏
Three-periodic nets and tilings: minimal nets
err2004-10-26
err0
errOAAI
errCharlotte Bonneau; Olaf Delgado-Friedrichs; Michael O'Keeffe; Omar M. Yaghi
err分享
err收藏
Euploidy in somatic cells from R6/2 transgenic Huntington's disease mice
err2005-09-13
err0
errOAAI
errÅsa Petersén; Ylva Stewénius; Maria Björkqvist; David Gisselsson
err分享
err收藏
err分享
err收藏
err分享
err收藏
The Poincaré-sphere approach to polarization: Formalism and new labs with Poincaré beams
err2016-11-01
err0
PREAI
errJoshua A. Jones; Anthony J. D’Addario; Brett L. Rojec; G. Milione; Enrique J. Galvez
err分享
err收藏
Efficient passage ranking for document databases
err1999-10-01
err42
errOAAI
errKaszkiel, M; Zobel, J; Sacks-Davis, R
err分享
err收藏
学者 查看更多内容