arrow
Return

Adversarial examples for extreme multilabel text classification

delete2022-11-04
delete3
delete
OA
AI
M
Mohammadreza Qaraei
R
Rohit Babbar *
DOI:10.1007/s10994-022-06263-zdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Extreme Multilabel Text Classification (XMTC) is a text classification problem in which, (i) the output space is extremely large, (ii) each data point may have multiple positive labels, and (iii) the data follows a strongly imbalanced distribution. With applications in recommendation systems and automatic tagging of web-scale documents, the research on XMTC has been focused on improving prediction accuracy and dealing with imbalanced data. However, the robustness of deep learning based XMTC models against adversarial examples has been largely underexplored. In this paper, we investigate the behaviour of XMTC models under adversarial attacks. To this end, first, we define adversarial attacks in multilabel text classification problems. We categorize attacking multilabel text classifiers as (a) positive-to-negative, where the target positive label should fall out of top-k predicted labels, and (b) negative-to-positive, where the target negative label should be among the top-k predicted labels. Then, by experiments on APLC-XLNet and AttentionXML, we show that XMTC models are highly vulnerable to positive-to-negative attacks but more robust to negative-to-positive ones. Furthermore, our experiments show that the success rate of positive-to-negative adversarial attacks has an imbalanced distribution. More precisely, tail classes are highly vulnerable to adversarial attacks for which an attacker can generate adversarial samples with high similarity to the actual data-points. To overcome this problem, we explore the effect of rebalanced loss functions in XMTC where not only do they increase accuracy on tail classes, but they also improve the robustness of these classes against adversarial attacks. The code for our experiments is available at https://github.com/xmc-aalto/adv-xmtc.
Keywords:
Extreme classification
Adversarial attacks
Multilabel problems
Text classification
Data imbalance

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.7K
Citations:
3.4W

Organization

A
Aalto University
Scholars:
1.6W
Papers: 1.5W
Citations: 2.1W
Cited Papers

Cited Papers

Non-Circadian Expression Masking Clock-Driven Weak Transcription Rhythms in U2OS Cells
err2014-07-09
err0
errOAAI
errJulia Hoffmann; Laura Symul; Anton Shostak; Tamás Fischer; Felix Naef; Michael Brunner
errShare
errSave
Evaluation of the diagnostic accuracy of an affordable rapid diagnostic test for African Swine Fever antigen detection in Lao People’s Democratic Republic
err2020-12-01
err0
errOAAI
errNina Matsumoto; Jarunee Siengsanan-Lamont; Laurence J. Gleeson; Bounlom Douangngeun; Watthana Theppangna; Syseng Khounsy; Phouvong Phommachanh; Tariq Halasa; Russell D. Bush; Stuart D. Blacksell
errShare
errSave
Geological Structural Surface Evaluation Model Based on Unascertained Measure
err2019-12-09
err0
errOAAI
errKang Zhao; Qing Wang; Yajing Yan; Junqiang Wang; Kui Zhao; Shuai Cao; Yongjun Zhang
errShare
errSave
errShare
errSave
errShare
errSave
Developmental Influences on Adult Intelligence
err
IF0
err2005-02-24
err0
PREAI
errK. Warner Schaie
errShare
errSave
Data scarcity, robustness and extreme multi-label classification
err2019-03-18
err81
errOAAI
errBabbar, Rohit; Schoelkopf, Bernhard
errShare
errSave
researcher View more