arrow
返回

A document representation framework with interpretable features using pre-trained word embeddings

delete2019-11-25
delete4
PRE
AI
N
Narendra Babu Unnam *
DOI:10.1007/s41060-019-00200-5delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We propose an improved framework for document representation using word embeddings. The existing models represent the document as a position vector in the same word embedding space. As a result, they are unable to capture the multiple aspects as well as the broad context in the document. Also, due to their low representational power, existing approaches perform poorly at document classification. Furthermore, the document vectors obtained using such methods have uninterpretable features. In this paper, we propose an improved document representation framework which captures multiple aspects of the document with interpretable features. In this framework, a document is represented in a different feature space by representing each dimension with a potential feature word with relatively high discriminating power. A given document is modeled as the distances between the feature words and the document. To represent a document, we have proposed two criteria for the selection of potential feature words and a distance function to measure the distance between the feature word and the document. Experimental results on multiple datasets show that the proposed model consistently performs better at document classification over the baseline methods. The proposed approach is simple and represents the document with interpretable word features. Overall, the proposed model provides an alternative framework to represent the larger text units with word embeddings and provides the scope to develop new approaches to improve the performance of document representation and its applications.
Keyword:
Text mining
Feature engineering
Document representation
Document classification
Word embeddings
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
International Journal of Data Science and Analytics
IF:
2.8
论文数:
1.1K
被引数:
1.3K

机构

I
International Institute of Information Technology Hyderabad
学者数:
779
论文数: 685
被引数: 5
引用论文

引用论文

err分享
err收藏
Theoretical Study of AlnNn, GanNn, and InnNn(n= 4, 5, 6) Clusters
err2002-01-30
err0
PREAI
errAnil K. Kandalam; Miguel A. Blanco; Ravindra Pandey
err分享
err收藏
Lithium ion adsorption–desorption properties on spinel Li4Mn5O12 and pH-dependent ion-exchange model
err2015-03-01
err0
PREAI
errJiali Xiao; Xiaoyao Nie; Shuying Sun; Xingfu Song; Ping Li; Jianguo Yu
err分享
err收藏
The Value of Aggressive Therapy in the Hypertensive Patient with Azotemia
err1969-12-01
err0
errOAAI
errWILLIAM J. MROCZEK; MICHAEL DAVIDOV; LILLIAN GAVRILOVICH; FRANK A. FINNERTY
err分享
err收藏
Fundus Lesions in Malignant Hypertension
err1986-11-01
err0
PREAI
errSohan Singh Hayreh; Gary E. Servais; Prem Singh Virdi
err分享
err收藏
Precise Underwater Distance Measurement by Dual Acoustic Frequency Combs
err2019-08-05
err0
PREAI
errHanzhong Wu; Zhiwen Qian; Haoyun Zhang; Xinyang Xu; Bin Xue; Jingsheng Zhai
err分享
err收藏
Representation learning for very short texts using weighted word embedding aggregation
err2016-09-01
err125
errOAAI
errDe Boom, Cedric; Van Canneyt, Steven; Demeester, Thomas; Dhoedt, Bart
err分享
err收藏
学者 查看更多内容