Return
MoTIF: An end-to-end multimodal road traffic scene understanding foundation model
Z
H
C
Z
Y
D
J
S
Z
DOI:10.1016/j.commtr.2025.100227.png)
Abstract
En 中文
• This model breaks through the cross-modal coding and structured text training technology. • End-to-end multimodal foundation model training requires only video and structured text. • The fine-tuning technique based on low-rank matrices and prompt engineering has significantly improved the scene understanding ability. • This study constructs a video structured standard dataset for multimodal foundation models of road traffic scene understanding.
Keywords:
Road traffic
Scene understanding
Multimodal foundation model
Fine-tuning
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
14.5
Papers:
216
Citations:
915
