arrow
返回

Adversarial examples for extreme multilabel text classification

delete2022-11-04
delete3
delete
OA
AI
M
Mohammadreza Qaraei
R
Rohit Babbar *
DOI:10.1007/s10994-022-06263-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Extreme Multilabel Text Classification (XMTC) is a text classification problem in which, (i) the output space is extremely large, (ii) each data point may have multiple positive labels, and (iii) the data follows a strongly imbalanced distribution. With applications in recommendation systems and automatic tagging of web-scale documents, the research on XMTC has been focused on improving prediction accuracy and dealing with imbalanced data. However, the robustness of deep learning based XMTC models against adversarial examples has been largely underexplored. In this paper, we investigate the behaviour of XMTC models under adversarial attacks. To this end, first, we define adversarial attacks in multilabel text classification problems. We categorize attacking multilabel text classifiers as (a) positive-to-negative, where the target positive label should fall out of top-k predicted labels, and (b) negative-to-positive, where the target negative label should be among the top-k predicted labels. Then, by experiments on APLC-XLNet and AttentionXML, we show that XMTC models are highly vulnerable to positive-to-negative attacks but more robust to negative-to-positive ones. Furthermore, our experiments show that the success rate of positive-to-negative adversarial attacks has an imbalanced distribution. More precisely, tail classes are highly vulnerable to adversarial attacks for which an attacker can generate adversarial samples with high similarity to the actual data-points. To overcome this problem, we explore the effect of rebalanced loss functions in XMTC where not only do they increase accuracy on tail classes, but they also improve the robustness of these classes against adversarial attacks. The code for our experiments is available at https://github.com/xmc-aalto/adv-xmtc.
Keyword:
Extreme classification
Adversarial attacks
Multilabel problems
Text classification
Data imbalance

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

A
Aalto University
学者数:
1.6W
论文数: 1.5W
被引数: 2.1W
引用论文

引用论文

Non-Circadian Expression Masking Clock-Driven Weak Transcription Rhythms in U2OS Cells
err2014-07-09
err0
errOAAI
errJulia Hoffmann; Laura Symul; Anton Shostak; Tamás Fischer; Felix Naef; Michael Brunner
err分享
err收藏
Evaluation of the diagnostic accuracy of an affordable rapid diagnostic test for African Swine Fever antigen detection in Lao People’s Democratic Republic
err2020-12-01
err0
errOAAI
errNina Matsumoto; Jarunee Siengsanan-Lamont; Laurence J. Gleeson; Bounlom Douangngeun; Watthana Theppangna; Syseng Khounsy; Phouvong Phommachanh; Tariq Halasa; Russell D. Bush; Stuart D. Blacksell
err分享
err收藏
Geological Structural Surface Evaluation Model Based on Unascertained Measure
err2019-12-09
err0
errOAAI
errKang Zhao; Qing Wang; Yajing Yan; Junqiang Wang; Kui Zhao; Shuai Cao; Yongjun Zhang
err分享
err收藏
err分享
err收藏
err分享
err收藏
Developmental Influences on Adult Intelligence
err
IF0
err2005-02-24
err0
PREAI
errK. Warner Schaie
err分享
err收藏
Data scarcity, robustness and extreme multi-label classification
err2019-03-18
err81
errOAAI
errBabbar, Rohit; Schoelkopf, Bernhard
err分享
err收藏
学者 查看更多内容