arrow
返回

An 8-bit Single Perceptron Processing Unit for Tiny Machine Learning Applications

delete2023-01-01
delete7
delete
OA
AI
M
Marco Crepaldi *
M
Mirco Di Salvo
A
Andrea Merello
DOI:10.1109/ACCESS.2023.3327517delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We present a tiny MultiLayer Perceptron (MLP) accelerator named Single Perceptron Linear Vector Processor (SPLVP) that aims at extending the capabilities of limited resources MCUs, enabling inference time speedup and main CPU off-load. It is based on a single perceptron hardware unit, enhanced with an additional accumulation input and scaling features, that is sequentially scheduled to cover all the nodes of the network. The accelerator supports both linear and Rectified Linear Unit (ReLU) activation and its firmware can be generated from 8-bit tflite quantized models. We also present a complete design toolchain that encompasses supervised learning, compilation, assembly, simulation, and device programming. The hardware support for extra accumulation input and scaling, together with the processor memory partitioning, are the key features that enable significant speedups. By solving a toy recognition problem based on image data captured from an infra-red camera, measurements show that the execution speed of SPLVP at 80MHz outperforms an ARM Cortex-M4 STM32L476 microcontroller by a factor of 9.2 when the same ANN is translated to MCU code using the STM CubeMX-Ai converter at the same clock frequency. SPLVP is synthesized on a low-cost and gate-count Cyclone 10 LP FPGA resulting in an 18% logic and 77% memory occupation. The SPLVP assembly code can be directly converted into a VHDL description that directly hardcodes the ANN. The execution speed of an ANN model for Iris classification, fully synthesized, improves by a factor of 209 compared to firmware execution on the MCU. To verify the operation of SPLVP and its design framework, we have designed various tiny Machine Learning (ML) classifiers, for which we briefly discuss the obtained performance and the preprocessing techniques used. Across all these classifiers, the obtained speedup compared to the STM32 is 8.3-14.9 $\times $ .
Keyword:
Neural processing unit
multilayer perceptron
single perceptron linear vector processor
fully connected neural network
FPGA
compiler
design toolchain
MCU
tiny machine learning

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

I
istituto italiano di tecnologia - iit
学者数:
9.1K
论文数: 6.9K
被引数: 8
引用论文

引用论文

Optimized R peak detection algorithm for ultra low power ECG systems
err2011-11-01
err0
PREAI
errSachin Shrestha; Tom Torfs; Hyejung Kim; Refet Firat Yazicioglu; Inaki Romero; Dilpreet Buxi; Torfinn Berset; Marco Altini
err分享
err收藏
An online generalized eigenvalue version of Laplacian Eigenmaps for visual big data
err2016-01-01
err29
errOAAI
errMalik, Zeeshan Khawar; Hussain, Amir; Wu, Jonathan
err分享
err收藏
A survey of numerical linear algebra methods utilizing mixed-precision arithmetic
err2021-03-19
err80
PREAI
errAbdelfattah, Ahmad; Anzt, Hartwig; Boman, Erik G.; Carson, Erin; Cojean, Terry; Dongarra, Jack; Fox, Alyson; Gates, Mark; Higham, Nicholas J.; Li, Xiaoye S.; Loe, Jennifer; Luszczek, Piotr; Pranesh, Srikara; Rajamanickam, Siva; Ribizel, Tobias; Smith, Barry F.; Swirydowicz, Kasia; Thomas, Stephen; Tomov, Stanimire; Tsai, Yaohung M.; Yang, Ulrike Meier
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容