返回
Semi-Automatic Dataset Annotation Applied to Automatic Violent Message Detection
DOI:10.1109/ACCESS.2024.3361404.png)
摘要
En 中文
Annotated corpora are indispensable tools to train computational models in Artificial Intelligence and Natural Language Processing. However, manual annotation is a costly, arduous, and time-consuming task, especially when the annotation is semantically complex. To address the problem, this work applies a methodology for semi-automatic annotation of datasets based on the Human-in-the-Loop paradigm. The methodology supports the building of a resource, that benefits from a fine-grained annotation, to aid in the detection of Spanish violent messages sourced from social media (Twitter/X). After implementing the proposed methodology for semi-automatic violence annotation, a high quality resource was obtained (hereafter referred to as VILLANOS). The methodology consists of annotating the dataset incrementally, which delivers an increase in annotator efficiency, thereby validating the suitability of the proposal. Annotation time was reduced by 52% compared to manual annotation and performance, by training a model with the VILLANOS dataset, obtains an $F_{1}$ of 85.2%. These results demonstrate the efficiency and effectiveness of the methodology, evidencing its validity.
Keyword:
Annotations
Task analysis
Social networking (online)
Artificial intelligence
Human in the loop
Data models
Cultural differences
Natural language processing
violent language
hate speech detection
assisted annotation
dataset construction
human-in-the-loop
active learning
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Research on the Efficiency of Wireless Power Transfer System Based on Multi-Auxiliary Transmitting Coils基于多辅助发射线圈的无线电能传输系统效率研究

