arrow
Return

A document representation framework with interpretable features using pre-trained word embeddings

delete2019-11-25
delete4
PRE
AI
N
Narendra Babu Unnam *
DOI:10.1007/s41060-019-00200-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We propose an improved framework for document representation using word embeddings. The existing models represent the document as a position vector in the same word embedding space. As a result, they are unable to capture the multiple aspects as well as the broad context in the document. Also, due to their low representational power, existing approaches perform poorly at document classification. Furthermore, the document vectors obtained using such methods have uninterpretable features. In this paper, we propose an improved document representation framework which captures multiple aspects of the document with interpretable features. In this framework, a document is represented in a different feature space by representing each dimension with a potential feature word with relatively high discriminating power. A given document is modeled as the distances between the feature words and the document. To represent a document, we have proposed two criteria for the selection of potential feature words and a distance function to measure the distance between the feature word and the document. Experimental results on multiple datasets show that the proposed model consistently performs better at document classification over the baseline methods. The proposed approach is simple and represents the document with interpretable word features. Overall, the proposed model provides an alternative framework to represent the larger text units with word embeddings and provides the scope to develop new approaches to improve the performance of document representation and its applications.
Keywords:
Text mining
Feature engineering
Document representation
Document classification
Word embeddings
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
International Journal of Data Science and Analytics
IF:
2.8
Papers:
1.1K
Citations:
1.3K

Organization

Cited Papers

Cited Papers

Theoretical Study of AlnNn, GanNn, and InnNn(n= 4, 5, 6) Clusters
err2002-01-30
err0
PREAI
errAnil K. Kandalam; Miguel A. Blanco; Ravindra Pandey
errShare
errSave
Lithium ion adsorption–desorption properties on spinel Li4Mn5O12 and pH-dependent ion-exchange model
err2015-03-01
err0
PREAI
errJiali Xiao; Xiaoyao Nie; Shuying Sun; Xingfu Song; Ping Li; Jianguo Yu
errShare
errSave
The Value of Aggressive Therapy in the Hypertensive Patient with Azotemia
err1969-12-01
err0
errOAAI
errWILLIAM J. MROCZEK; MICHAEL DAVIDOV; LILLIAN GAVRILOVICH; FRANK A. FINNERTY
errShare
errSave
Fundus Lesions in Malignant Hypertension
err1986-11-01
err0
PREAI
errSohan Singh Hayreh; Gary E. Servais; Prem Singh Virdi
errShare
errSave
Precise Underwater Distance Measurement by Dual Acoustic Frequency Combs
err2019-08-05
err0
PREAI
errHanzhong Wu; Zhiwen Qian; Haoyun Zhang; Xinyang Xu; Bin Xue; Jingsheng Zhai
errShare
errSave
Representation learning for very short texts using weighted word embedding aggregation
err2016-09-01
err125
errOAAI
errDe Boom, Cedric; Van Canneyt, Steven; Demeester, Thomas; Dhoedt, Bart
errShare
errSave
researcher View more