arrow
Return

NMF-based approach to automatic term extraction

delete2022-08-01
delete8
PRE
AI
A
Aliya Nugumanova
D
Darkhan Akhmed-Zaki
М
Мадина Мансурова
Y
Yerzhan Baiburin
A
Almasbek Maulit *
DOI:10.1016/j.eswa.2022.117179delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This work describes automatic term extraction approach based on the combination of the probabilistic topic modelling (PTM) and non-negative matrix factorization (NMF). Topic modeling algorithms including NMF-based ones do not require expensive and time-consuming manual annotations for domain terms, but only a corpus of domain documents. The topics emerge from the corpus documents without any supervision as sets of most probable words. This work is aimed to investigate how fully and precisely these most probable words from topics can reflect domain terminology. We run a series of experiments on the novel, qualitatively annotated dataset ACTER that was first used in the TermEval 2020 Shared Task. We compare five different NMF algorithms and four different NMF initializations when changing the number of topics extracted from documents and the number of most probable words extracted from topics in order to determine optimal combinations for best performance of term extraction. Finally, we compare the obtained optimal combinations of NMF with the competitive methods in TermEval 2020 and prove that our approach is second only to two much more sophisticated, domaindependent supervised methods.
Keywords:
Automatic term extraction
Probabilistic topic modeling
NMF
Unsupervised term extraction
ACTER dataset
TermEval shared task

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

A
Astana IT University
Scholars:
245
Papers: 131
Citations: 1
A
Al-Farabi Kazakh National University
Scholars:
3.0K
Papers: 1.5K
Citations: 1.6K