arrow
Return

Contextual stroke embedding for Chinese sequence labeling

delete2026-07-01
delete0
PRE
AI
Z
Zhang, Jian
W
Wenfan Chen
L
Li, Qi
W
Weiguo Sheng *
DOI:10.1016/j.csl.2026.102025delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sequence labeling is a fundamental while challenging task in natural language processing. In this paper, we propose a language model with a stroke-based scheme, called contextual stroke embedding (CSE), for Chinese sequence labeling. The proposed scheme, which is based on the inner structure of Chinese characters, aims to extract rich internal semantic information from Chinese stroke sequence as well as contextual syntactic information among Chinese characters. The scheme works by first decomposing Chinese character into stroke sequence. The resulting sequences are then used to train a stroke-level language model, which contains a two-layer BiLSTM network architecture to produce character representations. The resulting method has been evaluated on various Chinese sequence labeling tasks, including named entity recognition, Chinese word segmentation and part-of-speech, and compared with related methods. The results confirm the significance of the devised CSE. Specifically, the CSE is able to capture Chinese semantic structural information, thus greatly improving the performance of proposed method on various Chinese sequence labeling tasks. Further, the results also show that our method could achieve the best performance among non-BERT methods for Chinese sequence labeling and work well with BERT.
Keywords:
Deep learning
Nature language processing
Sequence labeling
Language model

Journal

C
Computer Speech and Language
IF:
3.4
Papers:
1.5K
Citations:
2.6K

Organization

H
Hangzhou Normal University
Scholars:
531
Papers: 147
Citations: 0
Cited Papers

Cited Papers

No cited papers available