返回
Knowledge Distillation With Multi-Objective Divergence Learning
DOI:10.1109/LSP.2021.3077414.png)
摘要
En 中文
Knowledge distillation has proven to be an effective model compression method that exploits the knowledge from a teacher model for supervising a student model by minimizing the distribution difference between the knowledge and the prediction produced by the student model. In this work, we focus on improving its performance from the perspective of divergence measures. A general form representing a family of divergence measures is introduced and a novel learning paradigm that jointly optimizes multiple measures is proposed by formalizing it as a multi-objective learning problem. Conditioned on Pareto optimality, the weights of different divergences are tuned in an automated way during training. Extensive experiments on multiple datasets show the proposed method can significantly improve the performance of student networks compared to the state-of-the-art methods for knowledge distillation. Codes are available at https://github.com/CMLOO/MoDiv.
Keyword:
Knowledge distillation
multi-objective optimazation
alpha-divergence
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.6
论文数:
1.1W
被引数:
1.7W
机构
引用论文
Introduction to Materials for Phytotaxonomy in the People's Republic of China《中华人民共和国植物分类学材料导论》
Taxon
IF0

