返回
A Context-Based Word Indexing Model for Document Summarization
DOI:10.1109/TKDE.2012.114.png)
摘要
En 中文
Existing models for document summarization mostly use the similarity between sentences in the document to extract the most salient sentences. The documents as well as the sentences are indexed using traditional term indexing measures, which do not take the context into consideration. Therefore, the sentence similarity values remain independent of the context. In this paper, we propose a context sensitive document indexing model based on the Bernoulli model of randomness. The Bernoulli model of randomness has been used to find the probability of the cooccurrences of two terms in a large corpus. A new approach using the lexical association between terms to give a context sensitive weight to the document terms has been proposed. The resulting indexing weights are used to compute the sentence similarity matrix. The proposed sentence similarity measure has been used with the baseline graph-based ranking models for sentence extraction. Experiments have been conducted over the benchmark DUC data sets and it has been shown that the proposed Bernoulli-based sentence similarity model provides consistent improvements over the baseline IntraLink and UniformLink methods [1].
Keyword:
Lexical association
text summarization
document indexing
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Comfort care in trauma patients without severe head injury: In-hospital complications as a trigger for goals of care discussions
Injury
IF0
An effective sentence-extraction technique using contextual information and statistical approaches for text summarization使用上下文信息和统计方法进行文本摘要的有效句子提取技术

