arrow
返回

Semantic spaces for improving language modeling

delete2014-01-01
delete17
PRE
AI
T
Tomáš Brychcín *
M
Miloslav Konopík
DOI:10.1016/j.csl.2013.05.001delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Language models are crucial for many tasks in NLP (Natural Language Processing) and n-grams are the best way to build them. Huge effort is being invested in improving n-gram language models. By introducing external information (morphology, syntax, partitioning into documents, etc.) into the models a significant improvement can be achieved. The models can however be improved with no external information and smoothing is an excellent example of such an improvement. In this article we show another way of improving the models that also requires no external information. We examine patterns that can be found in large corpora by building semantic spaces (HAL, COALS, BEAGLE and others described in this article). These semantic spaces have never been tested in language modeling before. Our method uses semantic spaces and clustering to build classes for a class-based language model. The class-based model is then coupled with a standard n-gram model to create a very effective language model. Our experiments show that our models reduce the perplexity and improve the accuracy of n-gram language models with no external information added. Training of our models is fully unsupervised. Our models are very effective for inflectional languages, which are particularly hard to model. We show results for five different semantic spaces with different settings and different number of classes. The perplexity tests are accompanied with machine translation tests that prove the ability of proposed models to improve performance of a real-world application. (C) 2013 Elsevier Ltd. All rights reserved.
Keyword:
Class-based language models
Semantic spaces
HAL
COALS
BEAGLE
Random Indexing
Purandare and Pedersen
Clustering
Inflectional languages
Machine translation

期刊

C
Computer Speech and Language
IF:
3.4
论文数:
1.5K
被引数:
2.6K

机构

U
university of west bohemia pilsen
学者数:
1.6K
论文数: 1.3K
被引数: 8
引用论文

引用论文

err分享
err收藏
Topic tracking language model for speech recognition
err2011-04-01
err23
PREAI
errWatanabe, Shinji; Iwata, Tomoharu; Hori, Takaaki; Sako, Atsushi; Ariki, Yasuo
err分享
err收藏
Follow‐up of 25 patients with treatable ataxia: A comprehensive case series study25例可治疗性共济失调患者的随访:一项综合性病例系列研究
err2022-04-20
err0
errOAAI
errMahmoud Reza Ashrafi; Elham Pourbakhtyaran; Mohammad Rohani; Bita Shalbafan; Ali Reza Tavasoli; Sareh Hosseinpour; Maryam Rasulinezhad; Zahra Rezaei; Ali Zare Dehnavi; Seyyed Mohammad Mahdi Hosseiny; Roya Haghighi; Homa Ghabeli; Morteza Heidari
err分享
err收藏
Morphology-based language modeling for conversational Arabic speech recognition
err2006-10-01
err73
errOAAI
errKirchhoff, Katrin; Vergyri, Dimitra; Bilmes, Jeff; Duh, Kevin; Stolcke, Andreas
err分享
err收藏
学者 查看更多内容