Return
LiLTv2: Language-substitutable Layout-image Transformer for Visual Information Extraction
DOI:10.1145/3708351.png)
Abstract
En 中文
Visual Information Extraction (VIE) has experienced substantial growth and heightened interest due to its pivotal role in intelligent document processing. However, most existing related pre-trained models typically can only process the data from a certain (set of) language(s)-often just English, representing a distinct limitation. To solve it, we present a Language-substitutable Layout-image Transformer (LiLTv2). It can be pre-trained just once on mono-lingual documents and then collaborate with off-the-shelf textual models in other languages during fine-tuning. Firstly, LiLTv2 utilizes a new dual-stream model architecture, one stream for substitutable text information and the other for layout and image information. Then, LiLTv2 has improved upon the optimization strategy and the diverse tasks adopted in the pre-training stage. Finally, we innovatively propose a teacher-student knowledge distillation learning with segment-level multi-modal features named SegKD. Extensive experimental results on widely used benchmarks can demonstrate the superior effectiveness of our method.
Keywords:
Visual Information Extraction
Multi-Modal Document Understanding
Self-Supervised Pre-Training
Journal
IF:
6
Papers:
2.0K
Citations:
5.4K

