arrow
返回

A Submodular Optimization Framework for Imbalanced Text Classification With Data Augmentation

delete2023-01-01
delete0
delete
OA
AI
E
Eyor Alemayehu
Y
Yi Fang *
DOI:10.1109/ACCESS.2023.3267669delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In the domain of text classification, imbalanced datasets are a common occurrence. The skewed distribution of the labels of these datasets poses a great challenge to the performance of text classifiers. One popular way to mitigate this challenge is to augment underwhelmingly represented labels with synthesized items. The synthesized items are generated by data augmentation methods that can typically generate an unbounded number of items. To select the synthesized items that maximize the performance of text classifiers, we introduce a novel method that selects items that jointly maximize the likelihood of the items belonging to their respective labels and the diversity of the selected items. Our proposed method formulates the joint maximization as a monotone submodular objective function, whose solution can be approximated by a tractable and efficient greedy algorithm. We evaluated our method on multiple real-world datasets with different data augmentation techniques and text classifiers and compared results with several baselines. The experimental results demonstrate the effectiveness and efficiency of the proposed method.
Keyword:
Data augmentation
Data models
Optimization
Predictive models
Perturbation methods
Task analysis
Greedy algorithms
text classification
imbalanced datasets
submodular optimization

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

S
Santa Clara University
学者数:
1.2K
论文数: 1.2K
被引数: 1.7K
引用论文

引用论文

err分享
err收藏
Hierarchical Data Augmentation and the Application in Text Classification
err2019-01-01
err19
errOAAI
errYu, Shujuan; Yang, Jie; Liu, Danlei; Li, Runqi; Zhang, Yun; Zhao, Shengmei
err分享
err收藏
err分享
err收藏
学者 查看更多内容