arrow
Return

Kitsune: Enabling Dataflow Execution on GPUs with Spatial

delete2025-12-01
delete0
PRE
AI
M
Michael Davies *
N
Neal Crago
K
Karthikeyan Sankaralingam
K
Keckle, Stephen
DOI:10.1145/3777466delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
State-of-the-art DL models are growing in size and complexity, with many modern models also increasing in heterogeneity of behavior. GPUs are still the dominant platform for DL applications, relying on a bulk-synchronous execution model which has many drawbacks and is ill-suited for the graph structure of DL applications. Many industry and academic works attempt to overcome these by employing vertical fusion-combining multiple sequential operations into a single kernel-but this approach still fails to realize three untapped opportunities: (1) the fact that many resources on the GPU are idle while only one operator executes due to temporal multiplexing of the SM; (2) lower energy from more intelligent on-chip data-movement which lends to higher performance in a power-provisioned environment. (3) inability to exploit reduction dimensions as a source of parallelism to ease pressure on batch size. This article explores relatively uncharted territory, answering the following key question: Can modest adjustments to the current GPU architecture enable efficient dataflow execution, thereby circumventing the constraints of vertical fusion without necessitating a clean-slate architecture design. We develop Kitsune-a set of primitives to construct spatial pipelines which enable dataflow execution on GPUs, and an end-to-end compiler based on PyTorch Dynamo. Across 5 challenge applications, Kitsune can provide up to 2.8x and 2.2x performance improvement as well as up to 99% and 45% off-chip traffic reduction for inference and training, respectively.
Keywords:
GPU architecture
spatial pipelining
artificial intelligence

Journal

A
ACM Transactions on Architecture and Code Optimization
IF:
1.8
Papers:
96
Citations:
1.1K

Organization

N
nvidia corporation
Scholars:
767
Papers: 439
Citations: 1
Cited Papers

Cited Papers

Survey of Machine Learning Accelerators
err2020-09-22
err0
errOAAI
errAlbert Reuther; Peter Michaleas; Michael Jones; Vijay Gadepally; Siddharth Samsi; Jeremy Kepner
errShare
errSave
Accelerating Deep Learning Inference with Cross-Layer Data Reuse on GPUs
err2020-01-01
err0
PREAI
errWang,Xueying; Li,Guangli; Dong,Xiao; Li,Jiansong; Liu,Lei; Feng,Xiaobing
errShare
errSave
TASO
err2019-10-27
err0
errOAAI
errZhihao Jia; Oded Padon; James Thomas; Todd Warszawski; Matei Zaharia; Alex Aiken
errShare
errSave
Compute Substrate for Software 2.0
err2021-03-01
err0
errOAAI
errJasmina Vasiljevic; Ljubisa Bajic; Davor Capalija; Stanislav Sokorac; Dragoljub Ignjatovic; Lejla Bajic; Milos Trajkovic; Ivan Hamer; Ivan Matosevic; Aleksandar Cejkov; Utku Aydonat; Tony Zhou; Syed Zohaib Gilani; Armond Paiva; Joseph Chu; Djordje Maksimovic; Stephen Alexander Chin; Zahi Moudallal; Akhmed Rakhmati; Sean Nijjar; Almeet Bhullar; Boris Drazic; Charles Lee; James Sun; Kei-Ming Kwong; James Connolly; Miles Dooley; Hassan Farooq; Joy Yu Ting Chen; Matthew Walker; Keivan Dabiri; Kyle Mabee; Rakesh Shaji Lal; Namal Rajatheva; Renjith Retnamma; Shripad Karodi; Daniel Rosen; Emilio Munoz; Andrew Lewycky; Aleksandar Knezevic; Raymond Kim; Allan Rui; Alexander Drouillard; David Thompson
errShare
errSave
errShare
errSave
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion
err2023-02-01
err0
PREAI
errSize Zheng; Siyuan Chen; Peidi Song; Renze Chen; Xiuhong Li; Shengen Yan; Dahua Lin; Jingwen Leng; Yun Liang
errShare
errSave
researcher View more