arrow
Return

Consistent multimodal pre-training for visual tokenization

delete2025-09-28
delete0
PRE
AI
T
Ting Pan
L
Lulu Tang
王新龙 (Xinlong Wang)
X
Xin Liu *
S
Shiguang Shan *
DOI:10.1007/s11432-024-4603-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal large language models (MLLMs) have recently demonstrated notable progress in understanding diverse visual context. Nevertheless, the overall performance of these large vision-language connecting models is highly related to a smaller vision-language pre-trained (CLIP) model at low resolution. Currently, this nesting vision-language alignment paradigm has hindered the development of a distinct vision foundation model for domain-specific multimodal tasks (e.g., OCR and document perception). In this paper, we explore a native high-resolution vision foundation model that is specifically designed for both image-level and region-level multimodal language tasks, clearly substituting the low-resolution CLIP models. Specifically, we introduce TAP-v2, a novel visual tokenizer that encodes general-purpose contextual information to enable comprehensive perception across diverse visual content.
Keywords:
foundation model
multimodal
representation learning
visual tokenization

Journal

Science China Information Sciences cover
Science China Information Sciences
IF:
7.6
Papers:
4.9K
Citations:
8.9K

Organization

I
Institute of Computing Technology
Scholars:
249
Papers: 113
Citations: 0