arrow
返回

TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation

delete2024-12-01
delete0
PRE
AI
R
Ruiping Liu
K
Kailun Yang *
A
Alina Roitberg
J
Jiaming Zhang
K
Kunyu Peng
Y
Yaonan Wang
R
Rainer Stiefelhagen
DOI:10.1109/TITS.2024.3455416delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Semantic segmentation benchmarks in the realm of autonomous driving are dominated by large pre-trained transformers, yet their widespread adoption is impeded by substantial computational costs and prolonged training durations. To lift this constraint, we look at efficient semantic segmentation from a perspective of comprehensive knowledge distillation and aim to bridge the gap between multi-source knowledge extractions and transformer-specific patch embeddings. We put forward the Transformer-based Knowledge Distillation (TransKD) framework which learns compact student transformers by distilling both feature maps and patch embeddings of large teacher transformers, bypassing the long pre-training process and reducing the FLOPs by ${>}85.0\%$ . Specifically, we propose two fundamental modules to realize feature map distillation and patch embedding distillation, respectively: 1) Cross Selective Fusion (CSF) enables knowledge transfer between cross-stage features via channel attention and feature map distillation within hierarchical transformers; 2) Patch Embedding Alignment (PEA) performs dimensional transformation within the patchifying process to facilitate the patch embedding distillation. Furthermore, we introduce two optimization modules to enhance the patch embedding distillation from different perspectives: 1) Global-Local Context Mixer (GL-Mixer) extracts both global and local information of a representative embedding; 2) Embedding Assistant (EA) acts as an embedding method to seamlessly bridge teacher and student models with the teacher's number of channels. Experiments on Cityscapes, ACDC, NYUv2, and Pascal VOC2012 datasets show that TransKD outperforms state-of-the-art distillation frameworks and rivals the time-consuming pre-training method.
Keyword:
Transformers
Semantic segmentation
Computer vision
Feature extraction
Knowledge engineering
Training
Computational modeling
Knowledge distillation
semantic segmentation
scene parsing
vision transformer
scene understanding

期刊

IEEE Transactions on Intelligent Transportation Systems 封面图
IEEE Transactions on Intelligent Transportation Systems
IF:
8.4
论文数:
9.5K
被引数:
6.3W

机构

K
karlsruhe institute of technology
学者数:
2.0W
论文数: 1.4W
被引数: 23
H
Helmholtz Association
学者数:
13.2W
论文数: 10.7W
被引数: 145
S
swiss federal institutes of technology domain
学者数:
9.0W
论文数: 8.0W
被引数: 163
H
hunan university
学者数:
4.5W
论文数: 3.3W
被引数: 70
学者 查看更多机构
引用论文

引用论文

The Infrared Spectra of First Transition Series Metal(II) Complexes of Picolinic AcidN-Oxide
err2006-12-05
err0
PREAI
errThomas P. E. Auf der Heyde; Craig S. Green; Alan T. Hutton; David A. Thornton
err分享
err收藏
err分享
err收藏
A Pilot Study of the Efficacy of the Unified Protocol for Transdiagnostic Treatment of Emotional Disorders in Treating Posttraumatic Psychopathology: A Randomized Controlled Trial
err2021-01-16
err0
errOAAI
errMeaghan L. O'Donnell; Winnie Lau; Katherine Chisholm; James Agathos; Jonathon Little; Sonia Terhaag; Rachel Brand; Andrea Putica; Alexander C. N. Holmes; Lynda Katona; Kim L. Felmingham; Kim Murray; Fardous Hosseiny; Matthew W. Gallagher
err分享
err收藏
OCNet: Object Context for Semantic SegmentationOCNet: 用于语义分割的对象上下文
err2021-05-24
err165
errOAAI
errYuan, Yuhui; Huang, Lang; Guo, Jianyuan; Zhang, Chao; Chen, Xilin; Wang, Jingdong
err分享
err收藏
Knowledge Distillation: A Survey知识蒸馏: 一项调查
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
err分享
err收藏
学者 查看更多内容