Return
Language-Driven Visual Data Generation for Zero-Shot HOI Detection
DOI:10.1109/TIP.2026.3705170.png)
Abstract
En 中文
Zero-shot human-object interaction (HOI) detection aims to recognize both seen and unseen interaction categories while detecting humans and objects in an image. However, due to the absence of training samples for unseen categories, existing methods often overfit on seen HOIs and struggle to generalize to unseen ones. To address this issue, we introduce a novel Language-Driven Visual Data Generation (LD-VDG) approach that generates pseudo visual features from textual semantics of unseen HOIs. This provides an innovative solution enabling generalization to unseen HOIs without relying on visual samples. Specifically, we first design a text-to-vision (T-V) adapter to align HOI text and visual features, trained on seen HOIs with paired image-text data. For unseen HOIs, we guide the large language model to produce multiple fine-grained textual descriptions based on HOI labels, which are then encoded by the vision-language model and transformed into pseudo visual features via the T-V adapter. After that, these pseudo features together with real features from seen HOIs are jointly used to train a transformer-based HOI detector. In this way, our method enables effective recognition of unseen HOIs by leveraging language-driven visual representations. Experimental results on standard datasets demonstrate that the proposed LD-VDG outperforms previous methods. In particular, it achieves superior performance on unseen categories under various zero-shot settings.
Keywords:
Human-object interaction detection
vision-language model
large language model
Journal
IF:
13.7
Papers:
1.0W
Citations:
8.4W

