arrow
返回

Enhancing Parameter Efficiency in Model Inference Using an Ultralight Inter-Transformer Linear Structure

delete2024-01-01
delete0
delete
OA
AI
H
Haoxiang Shi *
T
Tetsuya Sakai
DOI:10.1109/ACCESS.2024.3378518delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Pre-trained language models are the cornerstone of modern natural language processing and information retrieval. However, fine-tuning all the parameters reduces the efficiency of models both in training and inference owing to their increasingly heavy structures. Existing methods for parameter efficiency still require approximately 1 MB of storage and have approximately 10(7) operations during model deployment and inference. This puts a strain on the storage and processor capacity of end devices such as smartphones and IoT equipment, and slow model inference adversely affecting the user experience. To achieve more efficient and storage-friendly inference compared to mainstream methods, such as low-rank adaptation (LoRA) and Adapter, LayerConnect (hyper-network-assisted interlayer connectors) is proposed in this paper. Extensive experiments were conducted to validate the performance of LayerConnect for two essential tasks with completely different learning frameworks and purposes: natural language understanding (using the general language understanding evaluation (GLUE) benchmark) and information retrieval (using the a contextualized inverted list (COIL) framework). For both tasks, our LayerConnect saves up to 95.31% and 91.18% of parameters in LoRA and Adapter, respectively. In contrast, LayerConnect maintains performance degradation for GLUE and COIL to less than 8% and 3%, compared to LoRA. When compared to Adapter, the numbers become 5% and 3%, for GLUE and COIL, respectively. In addition, LayerConnect required approximately 100 kB of storage per task-specific trained model for both tasks and reduced the number of operations in the model inference by four orders of magnitude, reaching approximately 10(3).
Keyword:
Task analysis
Transformers
Adaptation models
Connectors
Tuning
Information retrieval
Training
Parameter efficiency
model inference
hypernetwork
pretrained language model
green AI

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

W
Waseda University
学者数:
1.0W
论文数: 8.7K
被引数: 8.3K
引用论文

引用论文

The adsorption of monolayer coatings on iron nanoparticles: Mössbauer spectroscopy and XANES results
err1998-11-01
err0
PREAI
errG Kataby; Yu Koltypin; J Rothe; J Hormes; I Felner; X Cao; A Gedanken
err分享
err收藏
Green AI绿色AI
err2020-11-17
err569
PREAI
errSchwartz, Roy; Dodge, Jesse; Smith, Noah A.; Etzioni, Oren
err分享
err收藏
err分享
err收藏
A Survey on Cross-Lingual Summarization
err2022-11-28
err26
errOAAI
errWang, Jiaan; Meng, Fandong; Zheng, Duo; Liang, Yunlong; Li, Zhixu; Qu, Jianfeng; Zhou, Jie
err分享
err收藏
err分享
err收藏
学者 查看更多内容