arrow
Return

Efficient convolution pooling on the GPU

delete2020-04-01
delete11
delete
OA
AI
S
Shunsuke Suita
T
Takahiro Nishimura
H
Hiroki Tokura
K
Koji Nakano *
Y
Yasuaki Ito
A
Akihiko Kasagi
T
Tsuguchika Tabaru
DOI:10.1016/j.jpdc.2019.12.006delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
The main contribution of this paper is to show efficient implementations of the convolution-pooling in the GPU, in which the pooling follows the multiple convolution. Since the multiple convolution and the pooling operations are performed alternately in earlier stages of many Convolutional Neural Networks (CNNs), it is very important to accelerate the convolution-pooling. Our new GPU implementation uses two techniques, (1) convolution interchange with direct sum, and (2) conversion to matrix multiplication. By these techniques, the computational and memory access cost are reduced. Further the convolution interchange is converted to matrix multiplication, which can be computed by cuBLAS very efficiently. Experimental results using Tesla V100 GPU show that our new GPU implementation compatible with cuDNN for the convolution-pooling is expected 2.90 times and 1.43 times faster for fp32 and fp16 than the multiple convolution and then the pooling by cuDNN, respectively. the most popular library of primitives to implement the CNNs in the GPU. (C) 2019 Elsevier Inc. All rights reserved.
Keywords:
Deep learning
Neural Networks
Convolution
Average pooling
GPU
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

F
fujitsu ltd
Scholars:
721
Papers: 512
Citations: 2
H
Hiroshima University
Scholars:
2.1W
Papers: 1.5W
Citations: 1.3W