返回
Latent semantics in language models
DOI:10.1016/j.csl.2015.01.004.png)
摘要
En 中文
This paper investigates three different sources of information and their integration into language modelling. Global semantics is modelled by Latent Dirichlet allocation and brings long range dependencies into language models. Word clusters given by semantic spaces enrich these language models with short range semantics. Finally, our own stemming algorithm is used to further enhance the performance of language modelling for inflectional languages. Our research shows that these three sources of information enrich each other and their combination dramatically improves language modelling. All investigated models are acquired in a fully unsupervised manner. We show the efficiency of our methods for several languages such as Czech, Slovenian, Slovak, Polish, Hungarian, and English, proving their multilingualism. The perplexity tests are accompanied by machine translation tests that prove the ability of the proposed models to improve the performance of a real-world application. (C) 2015 Elsevier Ltd. All rights reserved.
Keyword:
Language models
Latent Dirichlet allocation
Semantic spaces
Stemming
HAL
COALS
Random indexing
HPS
LDA
Machine translation
Moses
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
3.4
论文数:
1.5K
被引数:
2.6K
机构
引用论文
A unified language model for large vocabulary continuous speech recognition of Turkish土耳其语大词汇量连续语音识别的统一语言模型
SIGNAL PROCESSING
IF3.6
Representing word meaning and order information in a composite holographic lexicon
PSYCHOLOGICAL REVIEW
IF5.8

