Return
Exploring visual language models for driver gaze estimation: A task-based approach to debugging AI
DOI:10.1016/j.cviu.2025.104593.png)
Abstract
En 中文
• Evaluating complex VLMs for specific tasks like driver gaze estimation is challenging. • A proposed task-based methodology explores VLM knowledge/abilities for driver gaze estimation without exhaustive labelled data, aiding reliability. • VLMs struggle with fine-grained (e.g., 9-region) driver gaze classification in zero-shot scenarios, revealing limitations in spatial reasoning and interpreting subtle cues. • VLMs showed biases like prompt sensitivity, vocabulary choices, and use of image-relative coordinates that impact VLM gaze performance. • Understanding VLM limitations via exploration enables targeted strategies (prompting, fine-tuning) to significantly improve gaze estimation performance.
Keywords:
Driver monitoring systems
VLM
Explainability
Driver gaze estimation
Zero-shot evaluation
Open-vocabulary output of VLMs
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.5
Papers:
428
Citations:
7.3K

