Return
Consistent multimodal pre-training for visual tokenization
DOI:10.1007/s11432-024-4603-x.png)
Abstract
En 中文
Multimodal large language models (MLLMs) have recently demonstrated notable progress in understanding diverse visual context. Nevertheless, the overall performance of these large vision-language connecting models is highly related to a smaller vision-language pre-trained (CLIP) model at low resolution. Currently, this nesting vision-language alignment paradigm has hindered the development of a distinct vision foundation model for domain-specific multimodal tasks (e.g., OCR and document perception). In this paper, we explore a native high-resolution vision foundation model that is specifically designed for both image-level and region-level multimodal language tasks, clearly substituting the low-resolution CLIP models. Specifically, we introduce TAP-v2, a novel visual tokenizer that encodes general-purpose contextual information to enable comprehensive perception across diverse visual content.
Keywords:
foundation model
multimodal
representation learning
visual tokenization
Journal
IF:
7.6
Papers:
4.9K
Citations:
8.9K

