arrow
Return

Using Replicates in Information Retrieval Evaluation

delete2017-08-29
delete19
delete
OA
AI
E
Ellen M. Voorhees *
D
Daniel V. Samarov
I
Ian Soboroff
DOI:10.1145/3086701delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This article explores a method for more accurately estimating the main effect of the system in a typical testcollection-based evaluation of information retrieval systems, thus increasing the sensitivity of system comparisons. Randomly partitioning the test document collection allows for multiple tests of a given system and topic (replicates). Bootstrap ANOVA can use these replicates to extract system-topic interactions-something not possible without replicates-yielding a more precise value for the system effect and a narrower confidence interval around that value. Experiments using multiple TREC collections demonstrate that removing the topic-system interactions substantially reduces the confidence intervals around the system effect as well as increases the number of significant pairwise differences found. Further, the method is robust against small changes in the number of partitions used, against variability in the documents that constitute the partitions, and the measure of effectiveness used to quantify system effectiveness.
Keywords:
Information retrieval
statistical analysis
test collections
topic variance
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Information Systems cover
ACM Transactions on Information Systems
IF:
9.1
Papers:
1.2K
Citations:
4.7K

Organization

N
national institute of standards & technology (nist) - usa
Scholars:
9.7K
Papers: 9.0K
Citations: 4