返回
Efficient knowledge distillation using a shift window target-aware transformer
DOI:10.1007/s10489-024-06207-1.png)
摘要
En 中文
Target-aware Transformer (TaT) knowledge distillation effectively extracts information from intermediate layers but faces high computational costs for large feature maps. While the non-overlapping Patch-group distillation in TaT reduces complexity, it loses boundary information, affecting accuracy. We propose an improved Shifted Windows Target-aware Transformer (Swin TaT) knowledge distillation method, utilizing a hierarchical shift window strategy to preserve boundary information and balance computational efficiency. Our multi-scale approach optimizes Patch-group distillation with dynamic adjustment, ensuring effective local and global feature transfer. This flexible and efficient design enhances distillation performance, addressing previous limitations. The proposed Swin TaT method demonstrates exceptional performance across various architectures, with ResNet18 as the student network. It achieves 73.03% Top-1 accuracy on ImageNet1K, surpassing the SOTA by 1.06% while reducing parameters to approximately 46% less, and improves mIoU by 2.13% on COCOStuff10k.
Keyword:
Knowledge distillation
Target-aware Transformer (TaT)
Shifted window
Patch-group distillation
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Distant regulatory elements in a Sox10‐βGEO BAC transgene are required for expression of Sox10 in the enteric nervous system and other neural crest‐derived tissuesSox10-βgeo BAC转基因中的远距离调控元件是肠神经系统和其他神经源性组织中 Sox10 表达所必需的
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?

