返回
RefCap: image captioning with referent objects attributes
DOI:10.1038/s41598-023-48916-6.png)
摘要
En 中文
In recent years, significant progress has been made in visual-linguistic multi-modality research, leading to advancements in visual comprehension and its applications in computer vision tasks. One fundamental task in visual-linguistic understanding is image captioning, which involves generating human-understandable textual descriptions given an input image. This paper introduces a referring expression image captioning model that incorporates the supervision of interesting objects. Our model utilizes user-specified object keywords as a prefix to generate specific captions that are relevant to the target object. The model consists of three modules including: (i) visual grounding, (ii) referring object selection, and (iii) image captioning modules. To evaluate its performance, we conducted experiments on the RefCOCO and COCO captioning datasets. The experimental results demonstrate that our proposed method effectively generates meaningful captions aligned with users' specific interests.
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.9
论文数:
27.8W
被引数:
83.5W
机构
引用论文
Substrate specificity and inhibitor analyses of human steroid 5β-reductase (AKR1D1)人类类固醇5 β-还原酶 (AKR1D1) 的底物特异性和抑制剂分析
Steroids
IF0
Communication and reproductive behaviour in North American Jerusalem crickets (
Stenopelmatus
) (Orthoptera: Stenopelmatidae).北美耶路撒冷蟋蟀的交流和生殖行为 (
Stenopelmatus
) (直翅目: Stenopelmatidae)。
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
Enhancing the alignment between target words and corresponding frames for video captioning
PATTERN RECOGNITION
IF7.6

