arrow
返回

A Survey on Interpretability in Visual Recognition

delete2026-03-10
delete0
PRE
AI
Q
Qiyang Wan
R
Ruiping Wang
陈
陈熙霖 (Xilin Chen)
DOI:10.1109/TPAMI.2026.3672629delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous driving and medical diagnostics has accelerated the development of eXplainable AI (XAI). Distinct from generic XAI, visual recognition XAI is positioned at the intersection of vision and language, which represent the two most fundamental human modalities and form the cornerstones of multimodal intelligence. This paper provides a systematic survey of XAI in visual recognition by establishing a multi-dimensional taxonomy from a human-centered perspective based on intent, object, presentation, and methodology. Beyond categorization, we summarize critical evaluation desiderata and metrics, conducting an extensive qualitative assessment across different categories and demonstrating quantitative benchmarks within specific dimensions. Furthermore, we explore the interpretability of Multimodal Large Language Models and practical applications, identifying emerging trends and opportunities. By synthesizing these diverse perspectives, this survey provides an insightful roadmap to inspire future research on the interpretability of visual recognition models.
Keyword:
Explainable artificial intelligence
interpretability
visual recognition
XAI

期刊

IEEE Transactions on Pattern Analysis and Machine Intelligence 封面图
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
论文数:
1.0K
被引数:
9.8W

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

暂无论文信息