arrow
Return

Open CUDA convolution neural network inference implementation

delete2026-01-09
delete0
delete
OA
AI
P
Paulo A. C. Lopes *
DOI:10.1007/s10586-025-05895-9delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This work presents an open, efficient (fast) CUDA convolution neural network inference implementation specialized in some layers of popular nets like ResNet, VGG, and GoogLeNet. The proposed algorithm implements convolution directly instead of preprocessing with image to columns. Algorithm parameters are selected to meet constraints on global and shared memory access bandwidth, register usage, shared memory usage, and instructions per clock. Parallel arithmetic operations and memory access are achieved with several parallel blocks per streaming processor. Results comparing with state-of-the-art implementations are presented.
Keywords:
CNN
CUDA
GPU
Convolution
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Cluster Computing
IF:
0
Papers:
691
Citations:
1

Organization

I
instituto superior tecnico
Scholars:
106
Papers: 54
Citations: 0