1
Return

GUIDED++: Enhancing Discrimination with Conjunctive Verification for Fine-Grained Open-Vocabulary Object Detection

delete2026-08-09
delete0
PRE
AI
J
Jiaming Li
Z
Zhijia Liang
S
Shuangyin Liu
C
Chengpei Tang
G
Guanbin Li *
DOI:10.1007/s11263-026-02986-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Fine-grained open-vocabulary detection (FG-OVD) challenges models to detect objects described by complex, compositional language. The prevalent strategy of encoding descriptive phrases into a single, holistic vector creates significant semantic ambiguity, hindering true compositional understanding. This ambiguity manifests as two opposing failures: attribute over-representation, where a visually dominant attribute leads to mislocalization, and attribute under-representation, where the model fails to enforce a conjunctive match across all specified attributes, leading to false positives. In this paper, we propose GUIDED++, a framework that systematically resolves both issues through a dual strategy. First, we employ architectural decoupling to separate coarse-grained localization from fine-grained discrimination, effectively mitigating attribute over-representation. Second, we introduce a novel conjunctive multi-attribute verification mechanism to explicitly combat attribute under-representation. This mechanism decomposes a complex query into its atomic attribute conditions, computes a distinct similarity score for each, and uses a minimal aggregation function to enforce a strict logical ‘AND’, ensuring an object satisfies all required attributes. Extensive experiments show that GUIDED++ significantly outperforms existing methods and establishes a new state-of-the-art on challenging FG-OVD benchmarks, demonstrating a more robust approach to compositional visual reasoning.
Keywords:
Fine-grained open vocabulary object detection
Open vocabulary object detection
Vision language models
Conjunctive verification

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

C
College of Information Science and Technology
Scholars:
169
Papers: 76
Citations: 0
S
School of Computer Science and Engineering
Scholars:
1.1K
Papers: 511
Citations: 2
S
school of intelligent systems engineering
Scholars:
5
Papers: 2
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers