arrow
返回

Deep Learning Performance Characterization on GPUs for Various Quantization Frameworks

delete2023-10-18
delete3
delete
OA
AI
M
Muhammad Ali Shafique
A
Arslan Munir *
J
Joonho Kong
DOI:10.3390/ai4040047delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning is employed in many applications, such as computer vision, natural language processing, robotics, and recommender systems. Large and complex neural networks lead to high accuracy; however, they adversely affect many aspects of deep learning performance, such as training time, latency, throughput, energy consumption, and memory usage in the training and inference stages. To solve these challenges, various optimization techniques and frameworks have been developed for the efficient performance of deep learning models in the training and inference stages. Although optimization techniques such as quantization have been studied thoroughly in the past, less work has been done to study the performance of frameworks that provide quantization techniques. In this paper, we have used different performance metrics to study the performance of various quantization frameworks, including TensorFlow automatic mixed precision and TensorRT. These performance metrics include training time and memory utilization in the training stage along with latency and throughput for graphics processing units (GPUs) in the inference stage. We have applied the automatic mixed precision (AMP) technique during the training stage using the TensorFlow framework, while for inference we have utilized the TensorRT framework for the post-training quantization technique using the TensorFlow TensorRT (TF-TRT) application programming interface (API).We performed model profiling for different deep learning models, datasets, image sizes, and batch sizes for both the training and inference stages, the results of which can help developers and researchers to devise and deploy efficient deep learning models for GPUs.
Keyword:
optimization
deep learning
quantization
performance
TensorRT
automatic mixed precision

期刊

A
AI
IF:
5
论文数:
1.0K
被引数:
941

机构

K
Kansas State University
学者数:
9.5K
论文数: 8.1K
被引数: 1.3W
K
kyungpook national university (knu)
学者数:
1.8W
论文数: 1.8W
被引数: 14
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
SIRT1 Overexpression Antagonizes Cellular Senescence with Activated ERK/S6k1 Signaling in Human Diploid Fibroblasts
err2008-03-05
err0
errOAAI
errJing Huang; Qini Gan; Limin Han; Jian Li; Hai Zhang; Ying Sun; Zongyu Zhang; Tanjun Tong
err分享
err收藏
Artificial Intelligence and Data Fusion at the Edge
err2021-07-01
err55
PREAI
errMunir, Arslan; Blasch, Erik; Kwon, Jisu; Kong, Joonho; Aved, Alexander
err分享
err收藏
err分享
err收藏
学者 查看更多内容