arrow
返回

Unsupervised multi-sense language models for natural language processing tasks

delete2021-10-01
delete10
PRE
AI
J
Jihyeon Roh
S
Sungjin Park
B
Bo-Kyeong Kim
S
Sang-Hoon Oh
S
Soo-Young Lee *
DOI:10.1016/j.neunet.2021.05.023delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Existing language models (LMs) represent each word with only a single representation, which is unsuitable for processing words with multiple meanings. This issue has often been compounded by the lack of availability of large-scale data annotated with word meanings. In this paper, we propose a sense-aware framework that can process multi-sense word information without relying on annotated data. In contrast to the existing multi-sense representation models, which handle information in a restricted context, our framework provides context representations encoded without ignoring word order information or long-term dependency. The proposed framework consists of a context representation stage to encode the variable-size context, a sense-labeling stage that involves unsupervised clustering to infer a probable sense for a word in each context, and a multi-sense LM (MSLM) learning stage to learn the multi-sense representations. Particularly for the evaluation of MSLMs with different vocabulary sizes, we propose a new metric, i.e., unigram-normalized perplexity (PPLu), which is also understood as the negated mutual information between a word and its context information. Additionally, there is a theoretical verification of PPLu on the change of vocabulary size. Also, we adopt a method of estimating the number of senses, which does not require further hyperparameter search for an LM performance. For the LMs in our framework, both unidirectional and bidirectional architectures based on long short-term memory (LSTM) and Transformers are adopted. We conduct comprehensive experiments on three language modeling datasets to perform quantitative and qualitative comparisons of various LMs. Our MSLM outperforms single-sense LMs (SSLMs) with the same network architecture and parameters. It also shows better performance on several downstream natural language processing tasks in the General Language Understanding Evaluation (GLUE) and SuperGLUE benchmarks. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Language model
Neural language processing (NLP)
Multi-sense word modeling
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
7.8K
被引数:
3.0W

机构

M
Mokwon University
学者数:
184
论文数: 231
被引数: 122
引用论文

引用论文

Peroxisomal ABC Transporters: An Update
err2021-06-05
err0
errOAAI
errAli Tawbeh; Catherine Gondcaille; Doriane Trompier; Stéphane Savary
err分享
err收藏
Towards energy-autonomous wake-up receiver using Visible Light Communication
err2016-01-01
err0
errOAAI
errJoyce Sariol Ramos; Ilker Demirkol; Josep Paradells; Daniel Vossing; Karim M. Gad; Martin Kasemann
err分享
err收藏
Impurity effects on ionic-liquid-based supercapacitors
err2016-12-27
err0
errOAAI
errKun Liu; Cheng Lian; Douglas Henderson; Jianzhong Wu
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
Mesoscale physical variability affects zooplankton production in the Labrador Sea
err2009-05-01
err0
PREAI
errL. Yebra; R.P. Harris; E.J.H. Head; I. Yashayaev; L.R. Harris; A.G. Hirst
err分享
err收藏
Mathematical properties of optimal fluxes in cellular reaction networks at balanced growth
err2023-06-06
err0
errOAAI
errHugo Dourado; Wolfram Liebermeister; Oliver Ebenhöh; Martin J. Lercher
err分享
err收藏
学者 查看更多内容