1
Return

TrafficPerceiver: A multimodal large language model with reinforcement learning for unified challenging traffic scene perception

delete2026-03-31
delete0
PRE
AI
S
Senyun Kuang
Y
Y. Gao
S
Shijie Cong
Y
Yang Liu
危银涛 (Yintao Wei) *
DOI:10.26599/COMMTR.2026.9640008delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Understanding traffic scenes under diverse and challenging conditions is critical for intelligent transportation systems (ITSs). Existing methods primarily focus on ideal scenarios and often lack the ability to perform fine-grained perception or respond to human instructions. To address these limitations, we propose TrafficPerceiver, a unified multimodal framework based on a multimodal large language model (MLLM) that jointly supports both image understanding and target-oriented segmentation. To enhance the model's performance under adverse conditions such as rain, fog, and motion blur, we introduce a reinforcement learning (RL) optimization strategy based on group-relative policy optimization (GRPO), which encourages interpretable, instruction-following behavior. Additionally, we construct the challenging traffic scene understanding (CTSU) dataset, a large-scale dataset tailored to challenging traffic environments, with dense annotations for both segmentation and instruction-response tasks. Extensive experiments on both the DRAMA-ROLISP and CTSU datasets demonstrate that TrafficPerceiver achieves state-of-the-art performance in both understanding and segmentation tasks.
Keywords:
multimodal large language model (MLLM)
traffic scene perception
reinforcement learning (RL)
instruction-guided perception

Journal

Communications in Transportation Research cover
Communications in Transportation Research
IF:
14.5
Papers:
216
Citations:
915

Organization

T
tsinghua university
Scholars:
11.5W
Papers: 9.9W
Citations: 137
Cited Papers

Cited Papers

Citing Papers

Citing Papers