返回
Region-Focused Network for Dense Captioning
DOI:10.1145/3648370.png)
摘要
En 中文
Dense captioning is a very critical but under-explored task, which aims to densely detect localized regions-of-interest (RoIs) and describe them with natural language in a given image. Although recent studies tried to fuse multi-scale features from different visual instances to generate more accurate descriptions, their methods still suffer from the lack of exploration of relation semantic information in images, leading to less informative descriptions. Furthermore, indiscriminately fusing all visual instance features will introduce redundant information, resulting in poor matching between descriptions and corresponding regions. In thiswork, we propose a Region-Focused Network (RFN) to address these issues. Specifically, to fully comprehend the images, we first extract the object-level features, and encode the interaction and position relations between objects to enhance the object representations. Then, to decrease the interference fromredundant information about the target region, we extract the most relevant information to the region. Finally, a region-based Transformer is employed to compose and align the previous mined information and generate the corresponding descriptions. Extensive experiments on Visual Genome V1.0 and V1.2 datasets showthat our RFNmodel outperforms the state-of-the-art methods, thus verifying its effectiveness. Our code is available at https://github.com/VILAN-Lab/DesCap.
Keyword:
Dense captioning
interaction relation
region-focus
transformer
期刊
IF:
6
论文数:
2.0K
被引数:
5.4K
机构
引用论文
Cross-scale fusion detection with global attribute for dense captioning面向密集字幕的全局属性跨尺度融合检测
NEUROCOMPUTING
IF6.5
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue视觉对话中自适应视觉理解的学习双编码模型

