arrow
Return

Custom Hardware Inference Accelerator for TensorFlow Lite for Microcontrollers

delete2022-01-01
delete18
delete
OA
AI
E
Erez Manor
S
Shlomo Greenberg *
DOI:10.1109/ACCESS.2022.3189776delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, the need for the efficient deployment of Neural Networks (NN) on edge devices has been steadily increasing. However, the high computational demand required for Machine Learning (ML) inference on tiny microcontroller-based IoT devices avoids a direct software deployment on such resource-constrained edge devices. Therefore, various custom and application-specific NN hardware accelerators have been proposed to enable real-time Machine Learning (ML) inference on low-power and resource-limited edge devices. Efficient mapping of the computational load onto hardware and software resources is a key challenge for performance improvement while keeping low power and a low area footprint. High performance and yet low power embedded processors may be attained via the usage of hardware acceleration. This paper presents an efficient hardware-software framework to accelerate machine learning inference on edge devices using a modified TensorFlow Lite for Microcontroller (TFLM) model running on a Microcontroller (MCU) and a dedicated Neural Processing Unit (NPU) custom hardware accelerator, referred to as MCU-NPU. The proposed framework supports weight compression of pruned quantized NN models and exploits the pruned model sparsity to reduce computational complexity further. The proposed methodology has been evaluated by employing the MCU-NPU acceleration for various TFLM-based NN architectures using the common MLPerf Tiny benchmark. Experimental results demonstrate a significant speedup of up to 724x compared to a pure software implementation. For example, the resulting runtime for the CIFAR-10 classification is reduced from about 20 sec to only 37 ms using the proposed hardware acceleration. Moreover, the proposed hardware accelerator outperforms all the reference models optimized for edge devices in terms of inference runtime.
Keywords:
Computational modeling
Artificial neural networks
Hardware acceleration
Microcontrollers
Software
Kernel
Computational efficiency
TinyML
neural processing unit
TensorFlow-Lite for microcontrollers
hardware-software codesign

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

B
ben-gurion university of the negev
Scholars:
8.4K
Papers: 5.1K
Citations: 1
Cited Papers

Cited Papers

Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
err2019-05-30
err130
errOAAI
errWang, Erwei; Davis, James J.; Zhao, Ruizhe; Ng, Ho-Cheung; Niu, Xinyu; Luk, Wayne; Cheung, Peter Y. K.; Constantinides, George A.
errShare
errSave
Accelerating Deep Learning Inference in Constrained Embedded Devices Using Hardware Loops and a Dot Product Unit
err2020-01-01
err8
errOAAI
errVreca, Jure; Sturm, Karl J. X.; Gungl, Ernest; Merchant, Farhad; Bientinesi, Paolo; Leupers, Rainer; Brezocnik, Zmago
errShare
errSave
Hardware-Based Real-Time Deep Neural Network Lossless Weights Compression
err2020-01-01
err3
errOAAI
errMalach, Tomer; Greenberg, Shlomo; Haiut, Moshe
errShare
errSave
Growth and characterization of organometallic nonlinear optical TMTM single crystals
err2007-06-01
err0
PREAI
errK. Rajarajan; Preema C. Thomas; I. Vetha Potheher; Ginson P. Joseph; S.M. Ravi Kumar; S. Selvakumar; P. Sagayaraj
errShare
errSave
Preschool Phonological and Morphological Awareness As Longitudinal Predictors of Early Reading and Spelling Development in Greek
err2017-11-27
err0
errOAAI
errVassiliki Diamanti; Angeliki Mouzaki; Asimina Ralli; Faye Antoniou; Sofia Papaioannou; Athanassios Protopapas
errShare
errSave
An Overview of Machine Learning within Embedded and Mobile Devices-Optimizations and Applications
errSENSORS
IF3.5
err2021-06-28
err85
errOAAI
errAjani, Taiwo Samuel; Imoize, Agbotiname Lucky; Atayero, Aderemi A.
errShare
errSave
researcher View more