Return
Parameter-efficient online knowledge distillation for pretrained language models
DOI:10.1016/j.eswa.2024.126040.png)
Abstract
En 中文
With the advancement of natural language processing (NLP), the size of datasets and pre-trained language models (PLMs) has exponentially grown. These vast models exhibit robust capabilities in generation, comprehension, and multimodal processing. However, the complexity of parameters poses challenges for real-time and edge computing applications. Knowledge distillation (KD) has emerged as a common approach to address this issue by transferring prior knowledge from a large teacher to a smaller student model. Previous KD methods relied on offline full-parameter fine-tuning distillation, where the primary objective of the teacher model was to optimize downstream tasks, often neglecting the focus on KD itself. Such inefficient KD methods limit the potential of both teacher and student models. This study introduces a parameter-efficient online distillation method, PEKD, to tackle these challenges. In contrast to previous KD approaches, PEKD streamlines the distillation process into a single training stage and leverages parallel adapters to enhance fine-tuning efficiency. The teacher and student models, utilizing only 0.5% of the parameters on downstream tasks, achieve or surpass the accuracy of full fine-tuning methods. Extensive experiments on the GLUE benchmark dataset demonstrate the competitiveness of PEKD. The code for this study is accessible at https://github.com/Kawlyh/PEKD.
Keywords:
Online knowledge distillation
Pre-trained language model
Parameter compression
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
Organization
No organization information available

