arrow
Return

EcoFlow: Efficient Convolutional Dataflows on Low-Power Neural Network Accelerators

delete2024-09-01
delete1
PRE
AI
L
Lois Orosa *
S
Skanda Koppula
Y
Yaman Umuroglu
K
Konstantinos Kanellopoulos
J
Juan Gómez-Luna
M
Michaela Blott
K
Kees Vissers
O
Onur Mutlu
DOI:10.1109/TC.2023.3272282delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Dilated and transposed convolutions are widely used in modern convolutional neural networks (CNNs). These kernels are used extensively during CNN training and inference of applications such as image segmentation and high-resolution image generation. We find that commonly-used low-power CNN inference accelerators are not optimized for both these convolutional kernels. Dilated and transposed convolutions introduce significant zero padding when mapped to the underlying spatial architecture, significantly degrading performance and energy efficiency. Existing approaches that address this issue require significant design changes to the otherwise simple, efficient, and well-adopted architectures used to compute direct convolutions. To address this challenge, we propose EcoFlow, a new set of dataflows and mapping algorithms for dilated and transposed convolutions. These algorithms are tailored to execute efficiently on existing low-cost, small-scale spatial architectures and requires minimal changes to existing accelerators. At its core, EcoFlow eliminates zero padding through careful dataflow orchestration and data mapping tailored to the spatial architecture. We evaluate EcoFlow on CNN training workloads and Generative Adversarial Network (GAN) workloads. Experiments in SASiML, our new cycle-accurate simulator, show that, using a common CNN inference accelerator, EcoFlow 1) reduces end-to-end CNN training time between 7-85%, and 2) improves end-to-end GAN training performance between 29-42%, compared to state-of-the-art CNN dataflows.
Keywords:
Convolutional neural networks
Training
Computer architecture
Arrays
Kernel
Generative adversarial networks
Speech recognition
hardware accelerators

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

A
alphabet inc.
Scholars:
1.1K
Papers: 663
Citations: 0
S
swiss federal institutes of technology domain
Scholars:
9.0W
Papers: 8.0W
Citations: 163
D
deepmind
Scholars:
24
Papers: 11
Citations: 0
researcher View more organizations