arrow
返回

CREGEX: A Biomedical Text Classifier Based on Automatically Generated Regular Expressions

delete2020-01-01
delete7
delete
OA
AI
C
Christopher A. Flores
R
Rosa L. Figueroa *
J
Jorge E. Pezoa
Q
Qing Zeng‐Treitler
DOI:10.1109/ACCESS.2020.2972205delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
High accuracy text classifiers are used nowadays in organizing large amounts of biomedical information and supporting clinical decision-making processes. In medical informatics, regular expression-based classifiers have emerged as an alternative to traditional, discriminative classification algorithms due to their ability to model sequential patterns. This article presents CREGEX (Classifier Regular Expression), a biomedical text classifier based on an automatically generated regular-expressions-based feature space. We conceived an algorithm for automatically constructing an informative and discriminative regular-expressions-based feature space, suitable for binary and multiclass discrimination problems. Regular expressions are automatically generated from training texts using a coarse-to-fine text aligning method, which trades off the lexical variants of words, in terms of gender and grammatical number, and the generation of a feature space containing a large number of noisy features. CREGEX carries out feature selection by filtering keywords and also computes a confidence metric to classify test texts. Three de-identified datasets in Spanish, with information on smoking habits, obesity, and obesity types, were used here to assess the performance of CREGEX. For comparison, Support Vector Machine (SVM) and Na & x00EF;ve Bayes (NB) supervised classifiers were also trained with consecutive sequences of tokens (n-grams) as features. Results show that, in all the datasets used for evaluation, CREGEX not only outperformed both the SVM and NB classifiers in terms of accuracy and F-measure (p-value & x003C;0.05) but also used a fewer amount of training examples to achieve the same performance. Such a superior performance is attributed to the regular expressions; ability to represent complex text patterns.
Keyword:
Biomedical informatics
regular expressions
sequence alignment
text classification
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

G
George Washington University
学者数:
1.6W
论文数: 1.4W
被引数: 1.7W
U
universidad de concepcion
学者数:
8.4K
论文数: 6.4K
被引数: 8
引用论文

引用论文

Knowledge-enhanced document embeddings for text classification面向文本分类的知识增强文档嵌入
err2019-01-01
err96
errOAAI
errSinoara, Roberta A.; Camacho-Collados, Jose; Rossi, Rafael G.; Navigli, Roberto; Rezende, Solange O.
err分享
err收藏
err分享
err收藏
学者 查看更多内容