Return
Feature-based decoupled distillation
DOI:10.1016/j.knosys.2026.116658.png)
Abstract
En 中文
Feature-based knowledge distillation typically employs indiscriminate spatial alignment, implicitly assuming that all components within features are equally stable and transferable. In this work, we revisit this monolithic treatment through spectral analysis and reveal a clear feature heterogeneity across frequency bands: low-frequency components exhibit high cross-network consistency and resilience to perturbations, capturing stable, model-agnostic semantics, while high-frequency components show low similarity and sensitivity, encoding model-specific patterns and fine-grained details. Building on this finding, we propose Feature-based Decoupled Distillation, a divide-and-conquer framework that applies differentiated supervision to heterogeneous subspaces. Specifically, we introduce a content-adaptive decoupling module that leverages the teacher's global magnitude spectrum as a stable prior to separate dominant low-frequency components and complementary high-frequency components. We then impose rigid point-wise alignment on the low-frequency component to preserve semantic consistency, while adopting a relaxed energy-statistic matching objective for the high-frequency component to maintain the student's architectural flexibility. Extensive experiments on CIFAR-100, ImageNet, and MS-COCO validate the efficacy of our method and its strong generalization across architectures and tasks. Our code is available at: https://github.com/cloak-s/FDD.
Keywords:
Knowledge distillation
Feature decoupling
Frequency domain
Journal
K
IF:
7.6
Papers:
1.3W
Citations:
4.5W
Organization
Cited Papers
No cited papers available

