arrow
Return

LiLTv2: Language-substitutable Layout-image Transformer for Visual Information Extraction

delete2025-02-19
delete0
PRE
AI
J
Jiapeng Wang
Z
Zening Lin
D
Dayi Huang
L
Longfei Xiong
金连文 (Lianwen Jin)
DOI:10.1145/3708351delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual Information Extraction (VIE) has experienced substantial growth and heightened interest due to its pivotal role in intelligent document processing. However, most existing related pre-trained models typically can only process the data from a certain (set of) language(s)-often just English, representing a distinct limitation. To solve it, we present a Language-substitutable Layout-image Transformer (LiLTv2). It can be pre-trained just once on mono-lingual documents and then collaborate with off-the-shelf textual models in other languages during fine-tuning. Firstly, LiLTv2 utilizes a new dual-stream model architecture, one stream for substitutable text information and the other for layout and image information. Then, LiLTv2 has improved upon the optimization strategy and the diverse tasks adopted in the pre-training stage. Finally, we innovatively propose a teacher-student knowledge distillation learning with segment-level multi-modal features named SegKD. Extensive experimental results on widely used benchmarks can demonstrate the superior effectiveness of our method.
Keywords:
Visual Information Extraction
Multi-Modal Document Understanding
Self-Supervised Pre-Training

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

K
Kingsoft Off
Scholars:
2
Papers: 1
Citations: 0