Return
Short Text Topic Learning Using Heterogeneous Information Network
DOI:10.1109/TKDE.2022.3147766.png)
Abstract
En 中文
With the explosive growth of short texts on users' interests and preferences, learning discriminative and coherent latent topics from short texts is a critical and significative work, since many practical applications, such as e-commerce and recommendations, require semantic understandings that short texts convey explicitly and implicitly. However, existing short text topic learning methods face the challenge of fully capturing semantically related co-occurrence phrases. Therefore, this paper proposes a novel Heterogeneous Information Network-based Short Text Topic learning approach (HIN-ShoTT) in terms of parts of speech, without depending on any auxiliary information. Specifically, HIN-ShoTT can be decomposed into three phases: i) seeking semantic relations among words with different parts of speech, where HIN-ShoTT models multiple explicit and implicit semantic relations among words based on a Heterogeneous Information Network (HIN) in terms of parts of speech; ii) extracting co-occurrence phrases and filtering noises, where HIN-ShoTT defines parts-of-speech meta structures to guide co-occurrence phrase extraction and a self-adapting threshold filtering module is proposed for discarding noises; and iii) inferring topics, where HIN-ShoTT directly models the generative process of co-occurrence phrases to make topic learning effective with the abundant corpus-level information. Our experimental results on three real-world datasets not only show that HIN-ShoTT performs well, but also demonstrate that it is feasible to incorporate HIN into short text topic learning for accuracy improvement.
Keywords:
Semantics
Periodic structures
Optical wavelength conversion
Grammar
Peer-to-peer computing
Speech processing
Electronic mail
Short texts
topic learning
heterogeneous information network
parts of speech
meta structure
natural language processing
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W

