返回
A Memory-Efficient CNN Accelerator Using Segmented Logarithmic Quantization and Multi-Cluster Architecture
DOI:10.1109/TCSII.2020.3038897.png)
摘要
En 中文
This brief presents a memory-efficient CNN accelerator design for resource-constrained devices in Internet of Things (IoT) and autonomous systems. A segmented logarithmic (SegLog) quantization method is exploited to mitigate the on-chip memory and bandwidth requirements, thus accommodating more processing elements (PEs) in a given chip area to organize a reconfigurable multi-cluster architecture. The evaluation results show that SegLog quantization can achieve 6.4 x model compression with less than 2.5% accuracy loss on various CNNs. An ASIC implementation with 168 PEs configuration is validated in a 40-nm CMOS process, with 2.54 TOPs/W energy efficiency and 0.8 mm(2) chip area reported. The accelerator has also been implemented on FPGA with 1512 PEs configured and 468 kB on-chip memory, achieving a 1.29 GOPs/kB memory efficiency. Compared with the state-of-the-art accelerators, our ASIC implementation enhances area efficiency and arithmetic intensity by 1.94x and 5.62 x , while the FPGA implementation achieves the memory efficiency improvement by a factor of 2.34x.
Keyword:
Convolutional neural network (CNN)
dataflow
quantization
memory-efficient accelerator
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
4.9
论文数:
8.8K
被引数:
2.5W
机构
引用论文
Where Do I Belong? Prospective Relative Performance Information under High- and Low-Performing Reference Groups我属于哪里?高性能和低性能参考组下的预期相对性能信息
ACCOUNTING REVIEW
IF4.4
Synthesis Strategies and Nanoarchitectonics for High‐Performance Transition Metal Dichalcogenide Thin Film Field‐effect Transistors
ChemNanoMat
IF0
Factors associated with unplanned transfers among cancer patients at a freestanding acute rehabilitation facility
PM&R
IF0

