返回
Visual Probing: Cognitive Framework for Explaining Self-Supervised Image Representations
DOI:10.1109/ACCESS.2023.3242982.png)
摘要
En 中文
Recently introduced self-supervised methods for image representation learning provide on par or superior results to their fully supervised competitors, yet the corresponding efforts to explain the self-supervised approaches lag behind. Motivated by this observation, we introduce a novel visual probing framework for explaining the self-supervised models by leveraging probing tasks employed previously in natural language processing. The probing tasks require knowledge about semantic relationships between image parts. Hence, we propose a systematic approach to obtain analogs of natural language in vision, such as visual words, context, and taxonomy. Our proposal is grounded in Marr's computational theory of vision and concerns features like textures, shapes, and lines. We show the effectiveness and applicability of those analogs in the context of explaining self-supervised representations. Our key findings emphasize that relations between language and vision can serve as an effective yet intuitive tool for discovering how machine learning models work, independently of data modality. Our work opens a plethora of research pathways towards more explainable and transparent AI.
Keyword:
Computer vision
explainability
probing tasks self-supervised representation
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Quantifying spatial variability in shell midden formation in the Farasan Islands, Saudi Arabia量化沙特阿拉伯法拉桑群岛贝壳中层形成的空间变异性
PLOS ONE
IF0
Intercellular Adhesion Molecule 1 (ICAM-1) Gene Variant is Associated with Coronary Artery Calcification Independent of Soluble ICAM-1 Levels细胞间粘附分子1 (ICAM-1) 基因变异与冠状动脉钙化相关,与可溶性ICAM-1水平无关
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?

