arrow
Return

Semi-Supervised Image Captioning by Adversarially Propagating Labeled Data

delete2024-01-01
delete0
delete
OA
AI
D
Dong-Jin Kim *
T
Tae-Hyun Oh
J
Jinsoo Choi
I
In So Kweon
DOI:10.1109/ACCESS.2024.3423790delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present a novel data-efficient semi-supervised framework to improve the generalization of image captioning models. Constructing a large-scale labeled image captioning dataset is expensive in terms of labor, time, and cost. In contrast to manually annotating all the training samples, separately collecting uni-modal datasets is immensely easier, e.g., a large-scale image dataset and a sentence dataset. We leverage such massive unpaired image and caption data upon standard paired data by learning to associate them. To this end, our novel semi-supervised learning method assigns pseudo-labels to unpaired images and captions in an adversarial learning fashion, where the joint distribution of image and caption is learned. This approach shows noticeable performance improvement even in challenging scenarios, including out-of-task data and web-crawled data. We also show that our proposed method is theoretically well-motivated and has a favorable global optimal property. Our extensive and comprehensive empirical results on captioning datasets, followed by a comprehensive analysis of the scarcely-paired COCO dataset, demonstrate the consistent effectiveness of our method compared to competing ones.
Keywords:
Task analysis
Data models
Training
Semisupervised learning
Visualization
Natural languages
Generative adversarial networks
Closed captioning
Image captioning
unpaired captioning
semi-supervised learning
generative adversarial networks

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

H
hanyang university
Scholars:
2.9W
Papers: 2.7W
Citations: 36