arrow
返回

Optimizing Convolutions for Deep Learning Inference on ARM Cortex-M Processors

delete2024-08-01
delete1
delete
OA
AI
A
Antonio Maciá-Lillo
S
Sergio Barrachina
G
Germán Fabregat
M
Manuel F. Dolz *
DOI:10.1109/JIOT.2024.3395335delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We perform a series of optimizations on the convolution operator within the ARM common microcontroller units software interface standard for neural network (CMSIS-NN) library to improve the performance of deep learning tasks on Arduino development boards equipped with ARM Cortex-M4 and M7 microcontrollers. To this end, we develop custom microkernels that efficiently handle the internal computations required by the convolution operator via the lowering approach and the direct method, and we design two techniques to avoid register spilling. We also take advantage of all the RAM on the Arduino boards by reusing it as a scratchpad for the convolution filters. The integration of these techniques into CMSIS-NN, when invoked by TensorFlow Lite for microcontrollers for quantised versions of VGG, SqueezeNet, ResNet, and MobileNet-like convolutional neural networks enhances the overall inference speed by a factor ranging from 1.13xto 1.50x .
Keyword:
Program processors
Random access memory
Registers
Optimization
Convolution
Inference algorithms
Signal processing algorithms
ARM Cortex-M
common microcontroller units software interface standard for neural network (CMSIS-NN)
convolution operator
deep learning
edge computing
high performance
microcontrollers

期刊

IEEE Internet of Things Journal 封面图
IEEE Internet of Things Journal
IF:
8.9
论文数:
1.4W
被引数:
7.8W

机构

U
Universitat Jaume I
学者数:
4.7K
论文数: 4.8K
被引数: 6.1K
U
universitat d'alacant
学者数:
6.9K
论文数: 7.0K
被引数: 12
引用论文

引用论文

Enabling ImageNet-Scale Deep Learning on MCUs for Accurate and Efficient Inference
err2024-04-01
err1
errOAAI
errSadiq, Sulaiman; Hare, Jonathon; Craske, Simon; Maji, Partha; Merrett, Geoff
err分享
err收藏
err分享
err收藏
err分享
err收藏
Ethnicity and Immigration
err2009-10-30
err0
PREAI
errAndrew J. Fuligni; Diane L. Hughes; Niobe Way
err分享
err收藏
Reformulating the direct convolution for high-performance deep learning inference on ARM processors
err2023-02-01
err8
errOAAI
errBarrachina, Sergio; Castello, Adrian; Dolz, Manuel F.; Low, Tze Meng; Martinez, Hector; Quintana-Orti, Enrique S.; Sridhar, Upasana; Tomas, Andres E.
err分享
err收藏
学者 查看更多内容