arrow
返回

G-SLIDE: A GPU-Based Sub-Linear Deep Learning Engine via LSH Sparsification

delete2022-11-01
delete4
delete
OA
AI
Z
Zaifeng Pan
张峰 封面图
张峰 (Feng Zhang) *
H
Hourun Li
张晨阳 封面图
张晨阳 (Chenyang Zhang)
X
Xiaoyong Du
D
Dong Deng
DOI:10.1109/TPDS.2021.3132493delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning has been one of the trendiest research topics. However, as data quantities rise exponentially, training large neural networks can become prohibitively expensive with billions of parameters. Fortunately, recent research has discovered that not all of the computations in traditional network training are necessary. By selectively sparsifying the majority of the neurons during training, we can still obtain acceptable accuracy. SLIDE, a C++ OpenMP-based sub-linear deep learning engine, has been developed in this situation. SLIDE uses the algorithm of locality sensitive hashing (LSH) to query neurons with high activation in sub-linear time. It achieves a remarkable speedup in training large fully-connected networks by making use of the network sparsity as well as multi-core parallelism. However, SLIDE is limited to CPUs, ignoring the popular GPU devices with greater parallel potential and computational capability. In this article, we propose G-SLIDE, a GPU-based sub-linear deep learning engine, which combines the benefits of SLIDE's adaptive sparsification algorithms with GPUs' high performance. The main challenges in developing G-SLIDE are efficiently using LSH to sparsify networks and training the special sparse neural networks on the GPU. To address these challenges, we propose several novel solutions, such as specific data formats and appropriate workload partitioning for threads to fully utilize the GPU resources. We evaluate G-SLIDE on two extremely sparse datasets with a 2080 Ti GPU, and the results demonstrate that for the time of one training epoch, G-SLIDE can achieve more than 16.4x speedup over SLIDE on a 32-core/64-thread CPU. Furthermore, on the same platform, G-SLIDE can earn an average of 16.2x speedup over TensorFlow-GPU and 30.8x speedup over TensorFlow-CPU.
Keyword:
Graphics processing units
Training
Deep learning
Neurons
Biological neural networks
Engines
Message systems
GPU
machine learning system
adaptive sparsity
sparse neural network
LSH

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

R
Renmin University of China
学者数:
8.1K
论文数: 7.7K
被引数: 1.1W
R
rutgers university new brunswick
学者数:
2.3W
论文数: 1.9W
被引数: 32
引用论文

引用论文

err分享
err收藏
Ocrelizumab depletes T-lymphocytes more than rituximab in multiple sclerosis
err2021-04-01
err0
PREAI
errNicola Capasso; Agostino Nozzolillo; Giulia Scalia; Roberta Lanzillo; Antonio Carotenuto; Marcello De Angelis; Martina Petruzzo; Francesco Saccà; Cinzia Valeria Russo; Vincenzo Brescia Morra; Marcello Moccia
err分享
err收藏
err分享
err收藏
Automated Performance Modeling of HPC Applications Using Machine Learning
err2020-05-01
err23
PREAI
errSun, Jingwei; Sun, Guangzhong; Zhan, Shiyan; Zhang, Jiepeng; Chen, Yong
err分享
err收藏
Mössbauer spectroscopical investigation of the exchange biased Fe/MnF2 interface
err2006-10-27
err0
PREAI
errB. Sahoo; W. A. A. Macedo; W. Keune; V. Kuncser; J. Eisenmenger; J. Nogués; I. K. Schuller; I. Felner; Kai Liu; R. Röhlsberger
err分享
err收藏
err分享
err收藏
学者 查看更多内容