Return
Semantic-Spatial Guided Reasoning for Human-Object Interaction Detection
DOI:10.1109/LSP.2026.3652953.png)
Abstract
En 中文
Human-Object Interaction (HOI) detection requires not only recognizing <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">what</i> the interaction is but also understanding <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">where</i> it occurs. Although recent methods have achieved remarkable progress, they often lack effective joint modeling of spatial and semantic information, which is essential for accurate reasoning in complex scenes. In this paper, we propose a Semantic-Spatial Guided Reasoning (SSGR) framework that performs interaction reasoning by jointly modeling global semantic cues and fine-grained spatial priors. Specifically, SSGR constructs pair-specific spatial layouts to encode detailed spatial relationships and introduces a global semantic decoder to learn category-aware semantic representations. A semantic-spatial guided reasoning module further adaptively fuses these complementary cues, enabling unified reasoning and more discriminative interaction understanding. Extensive experiments on HICO-DET and V-COCO demonstrate that SSGR consistently outperforms prior methods under both standard and zero-shot settings, validating the effectiveness of our semantic-spatial reasoning paradigm.
Keywords:
Human-object interaction detection
semantic-spatial guidance
relation detection
Journal
I
IF:
3.9
Papers:
596
Citations:
0

