arrow
Return

Optimizing Standard Convolution for Diverse Precision on DCU

delete2025-11-01
delete0
PRE
AI
H
Haobo Hua
C
Chuangzheng Hou
W
Wen, Zhuxin
张祥凯 cover
张祥凯 (Xiangkai Zhang)
X
Xiaodong Yu
J
Jiandong Shang *
L
Litao Zhang
DOI:10.1007/s42514-025-00253-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Standard convolution remains a major performance bottleneck in modern deep neural networks. Although existing optimization libraries demonstrate effectiveness, they often underutilize key architectural features of emerging accelerators like DCUs, leading to suboptimal performance. To address this limitation, we propose a holistic, architecture-aware framework that systematically co-optimizes memory hierarchy and computational pipelines. The framework dynamically adapts to convolution parameters for maximal hardware utilization, with core contributions including: an innovative memory management strategy mitigating access conflicts, an adaptive computation pipeline balancing parallelism and data reuse, and a method bypassing API limitations to leverage underlying hardware instructions. On DCU hardware, our framework achieves significant speedups over MIOpen - delivering \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$3.09\times $$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$1.64\times $$\end{document} average acceleration for FP16 and FP32 precision respectively, while reducing end-to-end training time for ResNet and EfficientNet by \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$6.7\%$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$12.1\%$$\end{document}.
Keywords:
Convolution
Implicit GEMM
DCU
Performance optimization
Memory optimization

Journal

C
CCF Transactions on High Performance Computing
IF:
1.9
Papers:
38
Citations:
253

Organization

Z
Zhengzhou University
Scholars:
6.8W
Papers: 4.4W
Citations: 8.5W
Z
Zhengzhou University of Aeronautics
Scholars:
1.7K
Papers: 1.0K
Citations: 1.4K