arrow
Return

Parameter-efficient online knowledge distillation for pretrained language models

delete2025-03-01
delete0
PRE
AI
Y
Yukun Wang
王津 cover
王津 (Jin Wang) *
张学杰 (Xuejie Zhang)
DOI:10.1016/j.eswa.2024.126040delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the advancement of natural language processing (NLP), the size of datasets and pre-trained language models (PLMs) has exponentially grown. These vast models exhibit robust capabilities in generation, comprehension, and multimodal processing. However, the complexity of parameters poses challenges for real-time and edge computing applications. Knowledge distillation (KD) has emerged as a common approach to address this issue by transferring prior knowledge from a large teacher to a smaller student model. Previous KD methods relied on offline full-parameter fine-tuning distillation, where the primary objective of the teacher model was to optimize downstream tasks, often neglecting the focus on KD itself. Such inefficient KD methods limit the potential of both teacher and student models. This study introduces a parameter-efficient online distillation method, PEKD, to tackle these challenges. In contrast to previous KD approaches, PEKD streamlines the distillation process into a single training stage and leverages parallel adapters to enhance fine-tuning efficiency. The teacher and student models, utilizing only 0.5% of the parameters on downstream tasks, achieve or surpass the accuracy of full fine-tuning methods. Extensive experiments on the GLUE benchmark dataset demonstrate the competitiveness of PEKD. The code for this study is accessible at https://github.com/Kawlyh/PEKD.
Keywords:
Online knowledge distillation
Pre-trained language model
Parameter compression

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

No organization information available