arrow
返回

Region-Focused Network for Dense Captioning

delete2024-03-26
delete0
PRE
AI
Q
Qingbao Huang *
P
Pijian Li
Y
Youji Huang
F
Feng Shuang
蔡毅 封面图
蔡毅 (Yi Cai)
DOI:10.1145/3648370delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Dense captioning is a very critical but under-explored task, which aims to densely detect localized regions-of-interest (RoIs) and describe them with natural language in a given image. Although recent studies tried to fuse multi-scale features from different visual instances to generate more accurate descriptions, their methods still suffer from the lack of exploration of relation semantic information in images, leading to less informative descriptions. Furthermore, indiscriminately fusing all visual instance features will introduce redundant information, resulting in poor matching between descriptions and corresponding regions. In thiswork, we propose a Region-Focused Network (RFN) to address these issues. Specifically, to fully comprehend the images, we first extract the object-level features, and encode the interaction and position relations between objects to enhance the object representations. Then, to decrease the interference fromredundant information about the target region, we extract the most relevant information to the region. Finally, a region-based Transformer is employed to compose and align the previous mined information and generate the corresponding descriptions. Extensive experiments on Visual Genome V1.0 and V1.2 datasets showthat our RFNmodel outperforms the state-of-the-art methods, thus verifying its effectiveness. Our code is available at https://github.com/VILAN-Lab/DesCap.
Keyword:
Dense captioning
interaction relation
region-focus
transformer

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

G
guangxi university
学者数:
3.4W
论文数: 1.8W
被引数: 25
引用论文

引用论文

LSTM-based multi-label video event detection
err2017-12-18
err31
PREAI
errLiu, An-An; Shao, Zhuang; Wong, Yongkang; Li, Junnan; Su, Yu-Ting; Kankanhalli, Mohan
err分享
err收藏
Remote sensing of fish-processing in the Sundarbans Reserve Forest, Bangladesh: an insight into the modern slavery-environment nexus in the coastal fringe
err2020-09-17
err0
errOAAI
errBethany Jackson; Doreen S. Boyd; Christopher D. Ives; Jessica L. Decker Sparks; Giles M. Foody; Stuart Marsh; Kevin Bales
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容