Return
Prompt-based Weakly-supervised Vision-language Pre-training
DOI:10.1016/j.patrec.2025.06.020.png)
Abstract
En 中文
• PiTL uses weak cross-modal supervision, relying on LLM-generations of image labels. • PiTL mitigates overfitting with knowledge distillation and retrieval-augmented data. • PiTL unifies text and multi-modal encoders, and uses contrastive learning. • PiTL’s efficacy is evaluated across image-text retrieval, VE, VQA, and NLVR2 tasks. • PiTL’s components undergo a detailed analysis in retrieval tasks.
Keywords:
Weakly-supervised
Vision and language pre-training
Deep learning
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.3
Papers:
7.8K
Citations:
1.6W

