返回
Tens-embedding: A Tensor-based document embedding method
DOI:10.1016/j.eswa.2020.113770.png)
摘要
En 中文
A human is capable of understanding and classifying a text but a computer can understand the underlying semantics of a text when texts are represented in a way comprehensible by computers. The text representation is a fundamental stage in natural language processing (NLP). One of the main drawbacks of existing text representation approaches is that they only utilize one aspect or view of a text e.g. They only consider texts by their words while the topic information can be extracted from text as well. The term-document and document-topic matrix are two views of a text and contain complementary information. We use the strength of both views to extract a richer representation. In this paper, we propose three different text representation methods with the help of these two matrices and tensor factorization to utilize the power of both views. The proposed approach (Tens-Embedding) was applied in the tasks of text classification, sentence-level and document-level sentiment analysis and text clustering wherein the conducted experiments on 20 news groups, R52, R8, MR and IMDB datasets indicated the superiority of the proposed method in comparison with other document embedding techniques. (C) 2020 Elsevier Ltd. All rights reserved.
Keyword:
Natural language processing
Text classification
Text representation
Document embeddings
Tensor factorization
Topic modeling
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W

