返回
Efficient Model Switching in RRAM-Based DNN Accelerators
DOI:10.1109/TCAD.2025.3550403.png)
摘要
En 中文
Resistive random access memory (RRAM) has emerged as a promising technology for deep neural network (DNN) accelerators, but programming every weight in a DNN onto RRAM cells for inference can be both time-consuming and energy-intensive, especially when switching between different DNN models. This article introduces a hardware-aware multimodel merging (HA3M) framework designed to minimize the need for reprogramming by maximizing weight reuse, while taking into account the hardware constraints of the accelerator. The framework includes three key approaches: 1) crossbar (XB)-aware model mapping (XAMM); 2) block-based layer matching (BLM); and 3) multimodel retraining (MMR). XAMM reduces the XB usage of the preprogrammed model on RRAM XBs while preserving the model's structure. BLM reuses preprogrammed weights in a block-based manner, ensuring the inference process remains unchanged. MMR then equalizes the block-based matched weights across multiple models. Experimental results show that the proposed framework significantly reduces programming cycles in multi-DNN switching scenarios while maintaining or even enhancing accuracy, and eliminating the need for reprogramming.
Keyword:
Programming
Switches
Artificial neural networks
Computational modeling
Voltage
Hardware
Virtual machine monitors
Computer architecture
Integrated circuit modeling
Accuracy
Deep neural network (DNN)
model switching
multiple DNN application
resistive random access memory (RRAM)
RRAM programming
RRAM-based DNN accelerator
期刊
I
IF:
2.9
论文数:
668
被引数:
9.6K
机构
引用论文
Ultralow Power Neuromorphic Accelerator for Deep Learning Using Ni/HfO2/TiN Resistive Random Access Memory基于Ni/HfO₂/TiN阻变随机存取存储器的超低功耗神经形态加速器用于深度学习

