arrow
Return

Prompt-based Weakly-supervised Vision-language Pre-training

delete2025-07-09
delete0
delete
OA
AI
Z
Zixin Guo *
T
Tzu-Jui Julius Wang
S
Selen Pehlivan
A
Abduljalil Radman
M
Min Cao
J
Jorma Laaksonen
DOI:10.1016/j.patrec.2025.06.020delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• PiTL uses weak cross-modal supervision, relying on LLM-generations of image labels. • PiTL mitigates overfitting with knowledge distillation and retrieval-augmented data. • PiTL unifies text and multi-modal encoders, and uses contrastive learning. • PiTL’s efficacy is evaluated across image-text retrieval, VE, VQA, and NLVR2 tasks. • PiTL’s components undergo a detailed analysis in retrieval tasks.
Keywords:
Weakly-supervised
Vision and language pre-training
Deep learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

A
Aalto University
Scholars:
1.6W
Papers: 1.5W
Citations: 2.1W
Z
zenseact ab
Scholars:
1
Papers: 1
Citations: 0
V
VTT Technical Research Centre of Finland
Scholars:
124
Papers: 75
Citations: 0
S
soochow university
Scholars:
1.2W
Papers: 4.3K
Citations: 5
researcher View more organizations