arrow
返回

Fine-Tuning for Few-Shot Image Classification by Multimodal Prototype Regularization

delete2024-01-01
delete0
PRE
AI
J
Jiaxin Qi
张
张东 (Dong Zhang)
H
Hanwang Zhang
唐金辉 封面图
唐金辉 (Jinhui Tang) *
DOI:10.1109/TMM.2024.3379896delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large pre-trained vision-language models, such as CLIP [Radford et al. 2021], have demonstrated remarkable performance in few-shot image classification. To facilitate the rapid adaptation of CLIP in downstream tasks with limited visual samples, two primary frameworks have been proposed. The first framework centers on the image encoder and introduces a trainable visual classifier after the backbone to generate logits for each object class. Nevertheless, this framework heavily depends on limited visual features extracted by the pre-trained visual encoder, which can result in over-fitting issues. The second framework aims to optimize the text encoder by using trainable soft language prompts and computing logits for each class based on the similarity between image features and optimized prompt features. However, this framework encounters the issue of imperfect alignment between the representations extracted by the image and text encoders, making it difficult to fine-tune the language prompts using visual samples. This paper proposes a Multi-Modal Prototype Regularization (MMPR) method for CLIP-based few-shot fine-tuning for image classification. MMPR can address the challenges of effectively utilizing both image and text features. MMPR fine-tunes a classifier and regularizes its weights using both image-based (ImgPR) and text-based (TexPR) prototypes. ImgPR represents the mean of image representations within the same class, derived from the image encoder, to distill specific visual distribution knowledge for classifier adaptation. TexPR represents the hand-crafted prompt associated with the class, derived from the text encoder, to incorporate general encyclopedic knowledge and mitigate visual over-fitting. MMPR significantly leverages both image and text information without increasing computational complexity during the inference stage compared to existing methods. Experimental results on various challenging public benchmarks demonstrate the superiority of the proposed MMPR method over state-of-the-art methods.
Keyword:
Training
Visualization
Testing
Task analysis
Prototypes
Feature extraction
Tuning
Few-shot classification
large pre-trained vision-language models
model fine-tuning
prototype regularization

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

N
Nanyang Technological University
学者数:
4.9W
论文数: 4.8W
被引数: 8.1W
引用论文

引用论文

Interplanetary dust from the explosive dispersal of hydrated asteroids by impacts
err2003-05-01
err0
PREAI
errKazushige Tomeoka; Koji Kiriyama; Keiko Nakamura; Yasuhiro Yamahana; Toshimori Sekine
err分享
err收藏
Automated modal identification and tracking: Application to an iron arch bridge
err2016-02-28
err0
errOAAI
errAlessandro Cabboi; Filipe Magalhães; Carmelo Gentile; Álvaro Cunha
err分享
err收藏
Developmental brain changes during puberty and associations with mental health problems
err2023-04-01
err0
errOAAI
errNiousha Dehestani; Sarah Whittle; Nandita Vijayakumar; Timothy J. Silk
err分享
err收藏
High prevalence of malaria in a non-endemic setting among febrile episodes in travellers and migrants coming from endemic areas: a retrospective analysis of a 2013–2018 cohort
err2021-11-27
err0
errOAAI
errAlejandro Garcia-Ruiz de Morales; Covadonga Morcate; Elena Isaba-Ares; Ramon Perez-Tanoira; Jose A. Perez-Molina
err分享
err收藏
Powder metallurgical chromium
err1989-11-01
err0
PREAI
errR. Eck; H.P. Martinz; T. Sakaki; M. Kato
err分享
err收藏
Ingestion of corrosive acids
err1989-09-01
err0
PREAI
errShowkat Ali Zargar; Rakesh Kochhar; Birender Nagi; Saroj Mehta; Satish Kumar Mehta
err分享
err收藏
Elements of Quaternions四元数的元素
err
IF0
err2015-07-05
err0
PREAI
errWilliam Rowan Hamilton
err分享
err收藏
err分享
err收藏
学者 查看更多内容