arrow
返回

Adaptive Text Denoising Network for Image Caption Editing

delete2023-02-03
delete3
PRE
AI
M
Mengqi Yuan
B
Bing‐Kun Bao *
Z
Zhiyi Tan
徐
徐常胜 (Changsheng Xu)
DOI:10.1145/3532627delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image caption editing, which aims at editing the inaccurate descriptions of the images, is an interdisciplinary task of computer vision and natural language processing. As the task requires encoding the image and its corresponding inaccurate caption simultaneously and decoding to generate an accurate image caption, the encoder-decoder framework is widely adopted for image caption editing. However, existing methods mostly focus on the decoder, yet ignore a big challenge on the encoder: the semantic inconsistency between image and caption. To this end, we propose a novel Adaptive Text Denoising Network (ATD-Net) to filter out noises at the word level and improve the model's robustness at sentence level. Specifically, at the word level, we design a cross-attention mechanism called Textual Attention Mechanism (TAM), to differentiate the misdescriptive words. The TAM is designed to encode the inaccurate caption word by word based on the content of both image and caption. At the sentence level, in order to minimize the influence of misdescriptive words on the semantic of an entire caption, we introduce a Bidirectional Encoder to extract the correct semantic representation from the raw caption. The Bidirectional Encoder is able to model the global semantics of the raw caption, which enhances the robustness of the framework. We extensively evaluate our proposals on the MS-COCO image captioning dataset and prove the effectiveness of our method when compared with the state-of-the-arts.
Keyword:
Image caption editing
sequence editing
cross-modal semantic matching

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

err分享
err收藏
Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
err2020-05-01
err78
PREAI
errXu, Ning; Zhang, Hanwang; Liu, An-An; Nie, Weizhi; Su, Yuting; Nie, Jie; Zhang, Yongdong
err分享
err收藏
Remote sensing of fish-processing in the Sundarbans Reserve Forest, Bangladesh: an insight into the modern slavery-environment nexus in the coastal fringe
err2020-09-17
err0
errOAAI
errBethany Jackson; Doreen S. Boyd; Christopher D. Ives; Jessica L. Decker Sparks; Giles M. Foody; Stuart Marsh; Kevin Bales
err分享
err收藏
err分享
err收藏
Multilocus Metabarcoding of Terrestrial Leech Bloodmeal iDNA Increases Species Richness Uncovered in Surveys of Vertebrate Host Biodiversity
err2020-12-31
err0
PREAI
errMai Fahmy; Kalani M. Williams; Michael Tessler; Sarah R. Weiskopf; Evon Hekkala; Mark E. Siddall
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容