arrow
返回

Polysemy Deciphering Network for Robust Human-Object Interaction Detection

delete2021-04-19
delete32
delete
OA
AI
X
Xubin Zhong
丁
丁长兴 (Changxing Ding) *
X
Xian Qu
D
Dacheng Tao
DOI:10.1007/s11263-021-01458-8delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Human-Object Interaction (HOI) detection is important to human-centric scene understanding tasks. Existing works tend to assume that the same verb has similar visual characteristics in different HOI categories, an approach that ignores the diverse semantic meanings of the verb. To address this issue, in this paper, we propose a novel Polysemy Deciphering Network (PD-Net) that decodes the visual polysemy of verbs for HOI detection in three distinct ways. First, we refine features for HOI detection to be polysemy-aware through the use of two novel modules: namely, Language Prior-guided Channel Attention (LPCA) and Language Prior-based Feature Augmentation (LPFA). LPCA highlights important elements in human and object appearance features for each HOI category to be identified; moreover, LPFA augments human pose and spatial features for HOI detection using language priors, enabling the verb classifiers to receive language hints that reduce intra-class variation for the same verb. Second, we introduce a novel Polysemy-Aware Modal Fusion module, which guides PD-Net to make decisions based on feature types deemed more important according to the language priors. Third, we propose to relieve the verb polysemy problem through sharing verb classifiers for semantically similar HOI categories. Furthermore, to expedite research on the verb polysemy problem, we build a new benchmark dataset named HOI-VerbPolysemy (HOI-VP), which includes common verbs (predicates) that have diverse semantic meanings in the real world. Finally, through deciphering the visual polysemy of verbs, our approach is demonstrated to outperform state-of-the-art methods by significant margins on the HICO-DET, V-COCO, and HOI-VP databases. Code and data in this paper are available at .
Keyword:
Human-object interaction
Verb polysemy
Language priors
Attention model
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

International Journal of Computer Vision 封面图
International Journal of Computer Vision
IF:
9.3
论文数:
3.9K
被引数:
2.8W

机构

S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85
引用论文

引用论文

err分享
err收藏
Long-Term Longitudinal Patterns of Patient-Reported Fatigue After Breast Cancer: A Group-Based Trajectory Analysis
err2022-07-01
err0
errOAAI
errInes Vaz-Luis; Antonio Di Meglio; Julie Havas; Mayssam El-Mouhebb; Pietro Lapidari; Daniele Presti; Davide Soldato; Barbara Pistilli; Agnes Dumas; Gwenn Menvielle; Cecile Charles; Sibille Everhard; Anne-Laure Martin; Paul H. Cottu; Florence Lerebours; Charles Coutant; Sarah Dauchy; Suzette Delaloge; Nancy U. Lin; Patricia A. Ganz; Ann H. Partridge; Fabrice André; Stefan Michiels
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
A Data Mining Approach to Analyze Occupant Behavior Motivation
err2017-01-01
err0
errOAAI
errXinyuyang Ren; Yang Zhao; Wim Zeiler; Gert Boxem; Tingting Li
err分享
err收藏
学者 查看更多内容