arrow
Return

iMLBench: A Machine Learning Benchmark Suite for CPU-GPU Integrated Architectures

delete2021-07-01
delete11
PRE
AI
C
Chenyang Zhang
张峰 (Feng Zhang) *
X
Xiaoguang Guo
B
Bingsheng He
X
Xiao Zhang
X
Xiaoyong Du
DOI:10.1109/TPDS.2020.3046870delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Utilizing heterogeneous accelerators, especially GPUs, to accelerate machine learning tasks has shown to be a great success in recent years. GPUs bring huge performance improvements to machine learning and greatly promote the widespread adoption of machine learning. However, the discrete CPU-GPU architecture design with high PCIe transmission overhead decreases the GPU computing benefits in machine learning training tasks. To overcome such limitations, hardware vendors release CPU-GPU integrated architectures with shared unified memory. In this article, we design a benchmark suite for machine learning training on CPU-GPU integrated architectures, called iMLBench, covering a wide range of machine learning applications and kernels. We mainly explore two features on integrated architectures: 1) zero-copy, which means that the PCIe overhead has been eliminated for machine learning tasks and 2) co-running, which means that the CPU and the GPU co-run together to process a single machine learning task. Our experimental results on iMLBench show that the integrated architecture brings an average 7.1x performance improvement over the original implementations. Specifically, the zero-copy design brings 4.65x performance improvement, and co-running brings 1.78x improvement. Moreover, integrated architectures exhibit promising results from both performance-per-dollar and energy perspectives, achieving 6.50x performance-price ratio while 4.06x energy efficiency over discrete GPUs. The benchmark is open-sourced at https://github.com/ChenyangZhang-cs/iMLBench.
Keywords:
Computer architecture
Machine learning
Benchmark testing
Graphics processing units
Task analysis
Hardware
Training
Machine learning
benchmark
CPU
GPU
integrated architectures
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

R
Renmin University of China
Scholars:
8.1K
Papers: 7.7K
Citations: 1.1W
N
National University of Singapore
Scholars:
7.5W
Papers: 6.5W
Citations: 11.4W