返回
A network-based feature extraction model for imbalanced text data
DOI:10.1016/j.eswa.2022.116600.png)
摘要
En 中文
The explosive growth of text data has attracted many researchers to explore the efficient method to extract valuable hidden information. Many technologies, especially deep learning methods, have achieved great success in text analysis. However, the most powerful methods always require a considerable quantity of data for training, which may suffer from imbalanced data in some cases. In this paper, we propose a network-based Convolution Neural Network (NCNN) to mitigate the effect of imbalanced data. The proposed model first generates new synthetic samples for the imbalanced data based on the random walking of the network. Then an extra layer called Polar Layer is introduced to connect the output from the network model of the text to the classical CNN. Two electing strategies (n-NCNN and x-NCNN) are proposed to improve the performance of NCNN further. In the experimental section, the proposed model is applied to Reuters 21578 and WebKb. By comparing with six approaches, we prove the effectiveness of the proposed NCNN model on the imbalanced text data.
Keyword:
Complex Network
CNN
Text Analysis
Imbalanced Data
Random Walk
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0
RWO-Sampling: A random walk over-sampling approach to imbalanced data classificationRWO抽样: 一种非平衡数据分类的随机游走过抽样方法
INFORMATION FUSION
IF15.5

