返回
Collection selection for managed distributed document databases
DOI:10.1016/S0306-4573(03)00008-6.png)
摘要
En 中文
In a distributed document database system, a query is processed by passing it to a set of individual collections and collating the responses. For a system with many such collections, it is attractive to first identify a small subset of collections as likely to hold documents of interest before interrogating only this small subset in more detail. A method for choosing collections that has been widely investigated is the use of a selection index, which captures broad information about each collection and its documents. In this paper, we re-evaluate several techniques for collection selection. We have constructed new sets of test data that reflect one way in which distributed collections would be used in practice, in contrast to the more artificial division into collections reported in much previous work. Using these managed collections, collection ranking based on document surrogates is more effective than techniques such as CORI that are based on collection lexicons. Moreover, these experiments demonstrate that conclusions drawn from artificial collections are of questionable validity. (C) 2003 Elsevier Ltd. All rights reserved.
Keyword:
distributed document database
collection selection
meta-indexing
CORI
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
暂无机构信息
引用论文
IF0
Performance and genetic analysis of coast redwood cultivars for afforestation of converted grassland in California
New Forests
IF0
Applications of α-alkoxyorganocuprate reagents in the regiospecific synthesis of cyclic homoaldol products
Tetrahedron
IF0

