arrow
返回

Reduced precision floating-point optimization for Deep Neural Network On-Device Learning on microcontrollers

delete2023-12-01
delete7
delete
OA
AI
D
Davide Nadalini *
M
Manuele Rusci
L
Luca Benini
F
Francesco Conti
DOI:10.1016/j.future.2023.07.020delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Enabling On-Device Learning (ODL) for Ultra-Low-Power Micro-Controller Units (MCUs) is a key step for post-deployment adaptation and fine-tuning of Deep Neural Network (DNN) models in future TinyML applications. This paper tackles this challenge by introducing a novel reduced precision optimization technique for ODL primitives on MCU-class devices, leveraging the State-of-Art advancements in RISC-V RV32 architectures with support for vectorized 16-bit floating-point (FP16) Single-Instruction Multiple-Data (SIMD) operations. Our approach for the Forward and Backward steps of the Back Propagation training algorithm is composed of specialized shape transform operators and Matrix Multiplication (MM) kernels, accelerated with parallelization and loop unrolling. When evaluated on a single training step of a 2D Convolution layer, the SIMD-optimized FP16 primitives result up to 1.72x faster than the FP32 baseline on a RISC-V-based 8+1-core MCU. An average computing efficiency of 3.11 Multiply and Accumulate operations per clock cycle (MAC/clk) and 0.81 MAC/clk is measured for the end-to-end training tasks of a ResNet8 and a DS-CNN for Image Classification and Keyword Spotting, respectively - requiring 17.1 ms and 6.4 ms on the target platform to compute a training step on a single sample. Overall, our approach results more than two orders of magnitude faster than existing ODL software frameworks for single-core MCUs and outperforms by 1.6x previous FP32 parallel implementations on a Continual Learning setup.& COPY; 2023 Elsevier B.V. All rights reserved.
Keyword:
Parallel computing
Computer architecture
Open source software
Open architecture platforms
Deep learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.9K
被引数:
2.3W

机构

K
KU Leuven
学者数:
5.7W
论文数: 5.2W
被引数: 8.1W
U
University of Bologna
学者数:
4.5W
论文数: 3.8W
被引数: 4.1W
引用论文

引用论文

Vega: A Ten-Core SoC for IoT Endnodes With DNN Acceleration and Cognitive Wake-Up From MRAM-Based State-Retentive Sleep Mode
err2022-01-01
err48
errOAAI
errRossi, Davide; Conti, Francesco; Eggimann, Manuel; Di Mauro, Alfio; Tagliavini, Giuseppe; Mach, Stefan; Guermandi, Marco; Pullini, Antonio; Loi, Igor; Chen, Jie; Flamand, Eric; Benini, Luca
err分享
err收藏
A survey of transfer learning迁移学习研究综述
err2016-05-28
err0
errOAAI
errKarl Weiss; Taghi M. Khoshgoftaar; DingDing Wang
err分享
err收藏
DORY: Automatic End-to-End Deployment of Real-World DNNs on Low-Cost IoT MCUs
err2021-08-01
err76
errOAAI
errBurrello, Alessio; Garofalo, Angelo; Bruschi, Nazareno; Tagliavini, Giuseppe; Rossi, Davide; Conti, Francesco
err分享
err收藏
A Low-Power Transprecision Floating-Point Cluster for Efficient Near-Sensor Data Analytics用于高效近传感器数据分析的低功耗Transprecision浮点集群
err2022-05-01
err11
errOAAI
errMontagna, Fabio; Mach, Stefan; Benatti, Simone; Garofalo, Angelo; Ottavi, Gianmarco; Benini, Luca; Rossi, Davide; Tagliavini, Giuseppe
err分享
err收藏
Tests and calculations of structural element of temporary bridges
err2018-09-30
err0
PREAI
errAleksandr Ganyukov; Adil Kadyrov; Kyrmyzy Balabekova; Bakyt Kurmasheva
err分享
err收藏
512KiB RAM Is Enough! Live Camera Face Recognition DNN on MCU
err2019-10-01
err7
PREAI
errZemlyanikin, Maxim; Smorkalov, Alexander; Khanova, Tatiana; Petrovicheva, Anna; Serebryakov, Grigory
err分享
err收藏
Machine Learning at the Network Edge: A Survey
err2021-10-04
err167
errOAAI
errMurshed, M. G. Sarwar; Murphy, Christopher; Hou, Daqing; Khan, Nazar; Ananthanarayanan, Ganesh; Hussain, Faraz
err分享
err收藏
Robust Real-Time Embedded EMG Recognition Framework Using Temporal Convolutional Networks on a Multicore IoT Processor
err2020-04-01
err87
errOAAI
errZanghieri, Marcello; Benatti, Simone; Burrello, Alessio; Kartsch, Victor; Conti, Francesco; Benini, Luca
err分享
err收藏
学者 查看更多内容