返回
cuConv: CUDA implementation of convolution for CNN inference
DOI:10.1007/s10586-021-03494-y.png)
摘要
En 中文
Convolutions are the core operation of deep learning applications based on Convolutional Neural Networks (CNNs). Current GPU architectures are highly efficient for training and deploying deep CNNs, and are largely used in production. State-of-the-art implementations, however, present low efficiency for some commonly used network configurations. In this paper we propose a GPU-based implementation of the convolution operation for CNN inference that favors coalesced accesses, without requiring prior data transformations. Our experiments demonstrate that it yields notable performance improvements in a range of common CNN forward-propagation convolution configurations, with speedups of up to 2.29 x with respect to the best implementation in cuDNN, covering a relevant region in currently existing approaches. This improvement results in speedups of up to 7.4% for CNN online inference use cases.
Keyword:
Coalescing
Convolutional neural networks
cuDNN
Deep learning
GPU convolution
期刊
C
IF:
4.1
论文数:
5.1K
被引数:
7.5K
机构
引用论文
Identifying twins based on ocular region features using deep representations
APPLIED INTELLIGENCE
IF3.5
A novel and efficient xanthenic dye–organometallic ion‐pair complex for photoinitiating polymerization一种用于光引发聚合的新型高效的黄原胶染料-有机金属离子对配合物
Structural study of lanthanides(III) in aqueous nitrate and chloride solutions by EXAFS通过EXAFS对硝酸盐和氯化物水溶液中镧系元素 (III) 的结构研究
Gradient-based learning applied to document recognition基于梯度的学习在文档识别中的应用
PROCEEDINGS OF THE IEEE
IF25.9

