arrow
Return

Segmenting Brazilian legislative text using weak supervision and active learning

delete2024-09-26
delete0
PRE
AI
F
Felipe Siqueira *
D
Diany Pressato
F
Fabíola S. F. Pereira
N
Nádia Félix Felipe da Silva
E
Ellen Souza
M
Márcio S. Dias
A
André C. P. L. F. de Carvalho
DOI:10.1007/s10506-024-09419-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Legislative houses all over the world are adopting tools based on artificial intelligence to support their work. The incorporation of these tools can improve the analysis of the text of the proposed new laws and speed the preparation and discussion of new laws. The performance of artificial intelligence tools for text processing tasks is largely affected by the corpora used, which should ideally be adapted for the specific domain. When dealing with legislative corpora, text segmentation is often necessary due to the distinct purposes of legislative segments within the overall bill structure. While rule-based approaches can be effective in cases where the data follows a consistent format, they fail when inconsistencies arise in the formatting of legislative bills. In this study, we extensively investigate the use of weak supervision and active learning to accurately segment over 100,000 Brazilian federal legislative bills using a sequence tagging approach. The experiments demonstrated that both BERT and LSTM models achieved high statistical performance without the limitations of rule-based systems. In segmenting long documents beyond the limited context window of BERT, we find that simple moving windows suffice because the required context for accurate legislative segmentation is mostly local. We also conducted an analysis of transfer learning from our monolingual models to French, Italian, German, and English (US) legislative texts. According to our experimental results our models present non-trivial zero-shot and effective out-of-distribution fine-tuning performance, suggesting potential avenues for multilingual legislative segmentation without the need for computationally expensive models. The models, data, and code are publicly available at https://github.com/ulysses-camara/ulysses-segmenter.
Keywords:
Text segmentation
Legislative domain
Weak supervision
Active Learning
Portuguese data

Journal

Artificial Intelligence in Agriculture cover
Artificial Intelligence in Agriculture
IF:
12.4
Papers:
360
Citations:
1.7K

Organization

U
universidade de sao paulo
Scholars:
10.5W
Papers: 6.7W
Citations: 93