arrow
Return

Zone-YOLO: Vision-Language Object Detection Using Zone Prompt

delete2024-01-01
delete0
PRE
AI
J
Jiaxiong Yang
N
Ning Jia *
X
Xianhui Liu
R
Rui Fan
Y
Yougang Sun
W
Weidong Zhao
DOI:10.1109/TITS.2024.3510117delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Object detection in complex traffic scenarios is crucial for Intelligent Transportation Systems (ITS). At present, most real-time traffic object detection methods primarily rely on YOLO-style vision-only detectors, limiting their potential for further improvement. Vision-Language Object Detection (VLOD) has made promising progress currently, yet its adoption in the realm of ITS remains limited. Previous VLOD methods utilize text features in the classification task, without fully exploring their impact on the regression process for object localization. Besides, existing multi-modal fusion approaches fail to fuse text features with multi-scale image features at corresponding scales, which is detrimental to the representation capability of the model. In this work, we dive into the limitations above and introduce Zone-YOLO to improve the VLOD to a new level. Specifically, we propose Scale-Aware Modal Fusion (SAMF) to fully exploit the text and image features and learn to fuse the multi-modal representations seamlessly at different scales with channel-and modal-wise enhancement. Moreover, we present a novel Zone Prompt learning method to introduce text features into regression process and capture the zone-class-entity triple co-occurrence, which significantly improves the localization performance of the model. Extensive experiments show that Zone-YOLO outperforms the comparative methods by a considerable margin, achieving 55.1 AP, 72.1 AP(50) and 71.2 AP(L )on COCO. The competitive results on BDD100K and VisDrone2019 further demonstrate the superiority of Zone-YOLO on efficient traffic object detection.
Keywords:
Object detection
Feature extraction
Detectors
Visualization
Head
Fuses
Semantics
YOLO
Training
Encoding
Traffic object detection
vision-language model
multi-modal feature fusion
prompt learning

Journal

IEEE Transactions on Intelligent Transportation Systems cover
IEEE Transactions on Intelligent Transportation Systems
IF:
8.4
Papers:
9.5K
Citations:
6.3W

Organization

No organization information available