1
Return

MoTIF: An end-to-end multimodal road traffic scene understanding foundation model

delete2025-12-08
delete0
delete
OA
AI
Z
Zihe Wang
H
Haiyang Yu
C
Changxin Chen
Z
Zhiyong Cui *
Y
Yilong Ren
王子剑 cover
王子剑 (Zijian Wang)
D
Delan Kong
J
Jing Tian
S
Shoutong Yuan
Z
Zhiqiang Li
DOI:10.1016/j.commtr.2025.100227delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• This model breaks through the cross-modal coding and structured text training technology. • End-to-end multimodal foundation model training requires only video and structured text. • The fine-tuning technique based on low-rank matrices and prompt engineering has significantly improved the scene understanding ability. • This study constructs a video structured standard dataset for multimodal foundation models of road traffic scene understanding.
Keywords:
Road traffic
Scene understanding
Multimodal foundation model
Fine-tuning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Communications in Transportation Research cover
Communications in Transportation Research
IF:
14.5
Papers:
216
Citations:
915

Organization

B
Beihang University
Scholars:
5.0W
Papers: 4.0W
Citations: 37
L
Ltd.
Scholars:
1.1K
Papers: 593
Citations: 1
P
people's public security university of china
Scholars:
635
Papers: 401
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers