arrow
返回

Boosting Scene Graph Generation with Visual Relation Saliency

delete2023-01-05
delete7
PRE
AI
Y
Yong Zhang
Y
Yingwei Pan
T
Ting Yao *
R
Rui Huang
T
Tao Mei
C
Chang‐Wen Chen
DOI:10.1145/3514041delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The scene graph is a symbolic data structure that comprehensively describes the objects and visual relations in a visual scene, while ignoring the inherent perceptual saliency of each visual relation (i.e., relation saliency). However, humans often quickly allocate attention to important/salient visual relations in a scene. To align with such human perception of a scene, we explicitly model the perceptual saliency of visual relation in scene graph by upgrading each graph edge (i.e., visual relation) with an attribute of relation saliency. We present a new design, named as Saliency-guided Message Passing (SMP), that boosts the generation of such scene graph structure with the guidance from the visual relation saliency. Technically, an object interaction encoder is first utilized to strengthen object relation representations by jointly exploiting the appearance, semantic, and spatial relations in between. A branch is further leveraged to estimate the relation saliency of each visual relation by ordinal regression. Next, conditioned on the object and relation features (coupled with the estimated relation saliency), our SMP enhances scene graph generation by performing message passing over the objects and the most salient relations. Extensive experiments on VG-KR and VG150 datasets demonstrate the superiority of SMP for the scene graph generation. Moreover, we empirically validate the compelling generalizability of the learned scene graphs via SMP on downstream tasks like cross-model retrieval and image captioning.
Keyword:
Scene graph generation
relation saliency

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

H
hong kong polytechnic university
学者数:
3.0W
论文数: 4.1W
被引数: 921
C
Chinese University of Hong Kong
学者数:
3.4W
论文数: 3.2W
被引数: 5.6W
引用论文

引用论文

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Salient Object Detection: A Benchmark
err2015-12-01
err687
errOAAI
errBorji, Ali; Cheng, Ming-Ming; Jiang, Huaizu; Li, Jia
err分享
err收藏
Synthesis of n-type semiconducting diamond film using diphosphorus pentaoxide as the doping source以五氧化二磷为掺杂源合成n型半导体金刚石膜
err1990-10-01
err0
PREAI
errKen Okano; Hideo Kiyota; Tatsuya Iwasaki; Yoshitaka Nakamura; Yukio Akiba; Tateki Kurosu; Masamori Iida; Terutaro Nakamura
err分享
err收藏
Measurement of Talent in Team Handball: The Questionable Use of Motor and Physical Tests
err2005-01-01
err0
PREAI
errRonnie Lidor; Bareket Falk; Michal Arnon; Yoram Cohen; Gil Segal; Yael Lander
err分享
err收藏
学者 查看更多内容