arrow
返回

Active Learning for Biomedical Text Classification Based on Automatically Generated Regular Expressions

delete2021-01-01
delete20
delete
OA
AI
C
Christopher A. Flores *
R
Rosa L. Figueroa
J
Jorge E. Pezoa
DOI:10.1109/ACCESS.2021.3064000delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Biomedical text classification algorithms, which currently support clinical decision-making processes, call for expensive training texts due to the low availability of labeled corpus and the cost of manual annotation by specialized professionals. The active learning (AL) approach to classification heavily lessens such cost by reducing the number of labeled documents required to achieve specified performance. This article introduces a query strategy and a stopping criterion that transform CREGEX, a regular-expressions-based text classification algorithm, in an AL biomedical text classifier. The query strategy samples the training dataset, trading off the greedy learning achieved by the regular expressions classification precision and the conservative learning induced by text sequence alignment classification. The sustained reduction in the variance of the query strategy scores is used as a stopping criterion. The AL classifier was compared with Support Vector Machine (SVM), Naive Bayes (NB), and a classifier based on Bidirectional Encoder Representations from Transformers (BERT), using three datasets with biomedical information in Spanish on smoking habits, obesity, and obesity types. The learning curve results indicate that AL in CREGEX allowed to efficiently reduce the number of training examples for equal performance than the rest of the classifiers, obtaining areas under the learning curve greater than 85% in all cases. The stopping criterion applied to the AL process allowed to use, on average, approximately 32% to 50% of the total training examples with differences in performance concerning the maximum value of the learning curve not exceeding 2%. This performance demonstrates the effectiveness of using AL in a biomedical text classifier based on regular expressions, which is attributable to such expressions' ability to represent intricate sequential patterns in training texts considered most informative.
Keyword:
Training
Support vector machines
Obesity
Uncertainty
Bit error rate
Task analysis
Labeling
Active learning
regular expressions
natural language processing
text classification

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
universidad de concepcion
学者数:
8.4K
论文数: 6.4K
被引数: 8
引用论文

引用论文

Gaps in transitional care to adulthood for patients with cerebral palsy: a systematic review
err2023-08-08
err0
errOAAI
errDevon L. Mitchell; Nathan A. Shlobin; Emily Winterhalter; Sandi K Lam; Jeffrey S Raskin
err分享
err收藏
Optimization of WAAM Deposition Patterns for T-crossing FeaturesT型交叉特征的WAAM沉积模式优化
err2016-01-01
err0
errOAAI
errGiuseppe Venturini; Filippo Montevecchi; Antonio Scippa; Gianni Campatelli
err分享
err收藏
Age-specific Plasmodium parasite profile in pre and post ITN intervention period at a highland site in western Kenya
err2017-11-16
err0
errOAAI
errEdnah N. Ototo; Guofa Zhou; Lucy Kamau; Jenard P. Mbugi; Christine L. Wanjala; Maxwell Machani; Harrysone Atieli; Andrew K. Githeko; Guiyun Yan
err分享
err收藏
学者 查看更多内容