arrow
返回

Dynamic Contrastive Distillation for Image-Text Retrieval

delete2023-01-01
delete13
delete
OA
AI
J
Jun Rao
L
Liang Ding
S
Shuhan Qi *
M
Meng Fang
刘
刘扬 (Yang Liu)
沈力 封面图
沈力 (Li Shen)
D
Dacheng Tao
DOI:10.1109/TMM.2023.3236837delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The recent advancement in vision-and-language pretraining (VLP) has significantly improved the performance of cross-modal image-text retrieval (ITR) systems. However, the increasing size of VLP models presents a challenge for real-world deployment due to their high latency, making them unsuitable for practical search scenarios. To alleviate this problem, we present a novel plug-in dynamic contrastive distillation (DCD) framework to compress the large VLP models for the ITR task. Technically, we face the following two challenges: 1) the typical uni-modal metric learning approach is difficult to directly apply to cross-modal tasks due to the limited GPU memory to optimize too many negative samples during handling cross-modal fusion features. 2) it is inefficient to static optimize the student network from different hard samples, which affects distillation learning and student network optimization. We propose a method for multi-modal contrastive learning that balances training costs and effects. Our approach involves using a teacher network to identify hard samples for student networks to learn from, allowing the students to leverage the knowledge from pre-trained teachers and effectively learn from hard samples. To learn from hard sample pairs, we propose dynamic distillation to dynamically learn samples of different difficulties to balance better the difficulty of knowledge and students' self-learning ability. We successfully apply our proposed DCD strategy on two state-of-the-art vision-language pretrained models, i.e., ViLT and METER. Extensive experiments on MS-COCO and Flickr30K benchmarks show the effectiveness and efficiency of our DCD framework. We further provide in-depth analyses and discussions that explain how the performance improves.
Keyword:
Cross-modal retrieval
neural networks
contrastive learning

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
U
University of Liverpool
学者数:
2.8W
论文数: 2.5W
被引数: 3.5W
P
Peng Cheng Laboratory
学者数:
1.7K
论文数: 1.8K
被引数: 2.0K
学者 查看更多机构
引用论文

引用论文

Matryoshka Peek: Toward Learning Fine-Grained, Robust, Discriminative Features for Product Search
err2017-06-01
err3
PREAI
errKyaw, Zawlin; Qi, Shuhan; Gao, Ke; Zhang, Hanwang; Zhang, Luming; Xiao, Jun; Wang, Xuan; Chua, Tat-Seng
err分享
err收藏
err分享
err收藏
err分享
err收藏
Intra-Abdominal Splenosis Mimicking Metastatic Cancer
err2011-03-01
err0
PREAI
errNicholas J. Short; Teresa G. Hayes; Peeyush Bhargava
err分享
err收藏
Using collective system design to define and communicate organization goals and related solutions
err2017-10-12
err0
PREAI
errDavid Cochran; Gordon Schmidt; Jennifer Oxtoby; Mike Hensley; Jason Barnes
err分享
err收藏
Socio-economic status and outcomes for patients with age-related macular degeneration
errEye
IF0
err2019-03-11
err0
errOAAI
errPradnya More; Hussein Almuhtaseb; Dianna Smith; Simon Fraser; Andrew J. Lotery
err分享
err收藏
学者 查看更多内容