arrow
返回

A Deep Multi-Modal Explanation Model for Zero-Shot Learning

delete2020-01-01
delete16
delete
OA
AI
Y
Yu Liu *
T
Tinne Tuytelaars
DOI:10.1109/TIP.2020.2975980delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Zero-shot learning (ZSL) has attracted significant attention due to its capabilities of classifying new images from unseen classes. To perform the classification task for ZSL, learning visual and semantic embeddings has been the main research approach in existing literature. At the same time, generating complementary explanations to justify the classification decision has remained largely unexplored. In this paper, we propose to address a new and challenging task, namely explainable zero-shot learning (XZSL), which aims to generate visual and textual explanations to support the classification decision. To accomplish this task, we build a novel Deep Multi-modal Explanation (DME) model that incorporates a joint visual-attribute embedding module and a multi-channel explanation module in an end-to-end fashion. In contrast to existing ZSL approaches, our visual-attribute embedding is associated not only with the decision, but also with new visual and textual explanations. For visual explanations, we first capture several attribute activation maps (AAM) and then merge them into a class activation map (CAM) that visually infers which region of an image is relevant to the class. Textual explanations are generated from the multi-channel explanation module, jointly integrating three long short-term memory models (LSTMs) each of which is conditioned on a different feature representation. Additionally, we suggest that the DME model can retain explanatory consistency for similar instances and explanatory diversity for diverse instances. We conduct qualitative and quantitative experiments to assess the model for ZSL classification and explanation. Specifically, the ablation studies verify the effectiveness of the components in our model. Our results on three well-known datasets are competitive with prior approaches. More importantly, the joint training of our embedding and explanation modules demonstrates mutual performance improvements between ZSL classification and explanation. We shed more light on DME to analyze and diagnose its advantages and limitations.
Keyword:
Visualization
Semantics
Task analysis
Training
Image color analysis
Head
Extraterrestrial phenomena
Zero-shot learning
multi-modal explanation
visual-attribute embedding
class activation map
LSTM

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

K
KU Leuven
学者数:
5.7W
论文数: 5.2W
被引数: 8.1W
引用论文

引用论文

STUDIES OF ARTHROPOD‐BORNE VIRUS INFECTIONS IN QUEENSLAND
err1963-02-01
err0
PREAI
errRL Doherty; JG Carley; M Josephine Mackerras; Elizabeth N Marks
err分享
err收藏
Zero-Shot Learning With Transferred Samples
err2017-07-01
err84
PREAI
errGuo, Yuchen; Ding, Guiguang; Han, Jungong; Gao, Yue
err分享
err收藏
Zero-Shot Learning via Attribute Regression and Class Prototype Rectification
err2018-02-01
err51
PREAI
errLuo, Changzhi; Li, Zhetao; Huang, Kaizhu; Feng, Jiashi; Wang, Meng
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
The (European) Derisking State
err
IF0
err2023-05-17
err0
errOAAI
errDaniela Gabor
err分享
err收藏
学者 查看更多内容