返回
Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification
DOI:10.1109/TASLP.2024.3393722.png)
摘要
En 中文
Large-scale multi-label text classification (LMTC) aims at tagging each text with multiple relevant labels from a large label space, which typically demonstrates high sparsity, diversity, and skewness. To learn text representations in LMTC, a straightforward strategy is to learn a single vector to represent the whole text, yet limiting good generalization to diverse labels; another popular one is to learn specific representation per label via attention weighting, but excessively emphasizing tail labels restricts the overall performance. To cope with these limitations, we propose a novel LMTC framework, dubbed LADAR, which learns label-adaptive text representations to ensure high performance on large-scale labels. Specifically, we construct a representation pool for each text by collecting multi-layer features of the deep model as well as multi-granularity features of the text. Furthermore, all labels are adaptively matched to their most relevant representations to predict the final scores. Experiments over five benchmark datasets demonstrate the LADAR achieves highly superior results to state-of-the-art LMTC approaches. In particular, LADAR achieves significantly better performance on tail labels, e.g., 5.09% relative improvement on PSP@5 on the Amazon-670 K dataset than the best baseline.
Keyword:
Large-scale multi-label learning
text classification
long-tail issue
information extraction
期刊
I
IF:
5.1
论文数:
2.6K
被引数:
1.1W
机构
引用论文
Matrix Effects in the Detection of Pb and Ba in Soils Using Laser-Induced Breakdown Spectroscopy使用激光诱导击穿光谱检测土壤中Pb和Ba的基质效应
Definitive evidence that a single N-glycan among three glycans on inducible costimulator is required for proper protein trafficking and ligand binding明确的证据表明,在可诱导的共刺激物上的三个聚糖中,单个N-聚糖对于适当的蛋白质运输和配体结合是必需的

