arrow
Return

Effective and Robust Query-Based Stemming

delete2013-11-01
delete13
PRE
AI
J
Jiaul H. Paik *
S
Swapan K. Parui
D
Dipasree Pal
S
Stephen Robertson
DOI:10.1145/2536736.2536738delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Stemming is a widely used technique in information retrieval systems to address the vocabulary mismatch problem arising out of morphological phenomena. The major shortcoming of the commonly used stemmers is that they accept the morphological variants of the query words without considering their thematic coherence with the given query, which leads to poor performance. Moreover, for many queries, such approaches also produce retrieval performance that is poorer than no stemming, thereby degrading the robustness. The main goal of this article is to present corpus-based fully automatic stemming algorithms which address these issues. A set of experiments on six TREC collections and three other non-English collections containing news and web documents shows that the proposed query-based stemming algorithms consistently and significantly outperform four state of the art strong stemmers of completely varying principles. Our experiments also confirm that the robustness of the proposed query-based stemming algorithms are remarkably better than the existing strong baselines.
Keywords:
Algorithms
Experimentation
Performance
Corpus
stemming
suffix
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Information Systems cover
ACM Transactions on Information Systems
IF:
9.1
Papers:
1.2K
Citations:
4.7K

Organization

I
indian statistical institute kolkata
Scholars:
751
Papers: 814
Citations: 1
I
Indian Statistical Institute
Scholars:
1.7K
Papers: 1.8K
Citations: 1.2K