arrow
Return

Designing Programmable Accelerators for Sparse Tensor Algebra

delete2025-05-01
delete0
PRE
AI
K
Kalhan Koul
Z
Zhouhua Xie
M
Maxwell Strange
S
Sai Gautham Ravipati
B
Bo Cheng
O
Olivia Hsu
P
Po‐Han Chen
M
Mark Horowitz
F
Fredrik Kjølstad
P
Priyanka Raina
DOI:10.1109/MM.2025.3556611delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent research has focused on leveraging sparsity in hardware accelerators to improve the efficiency of applications spanning scientific computing to machine learning. Most such prior accelerators are fixed-function, which is insufficient for two reasons. First, applications typically include both dense and sparse components, and second, the algorithms that comprise these applications are constantly evolving. To address these challenges, we designed a programmable accelerator called Onyx for both sparse tensor algebra and dense workloads. Onyx extends a coarse-grained reconfigurable array (CGRA) optimized for dense applications with composable hardware primitives to support arbitrary sparse tensor algebra kernels. In this article, we show that we can further optimize Onyx by adding a small set of hardware features for parallelization that significantly increase both temporal and spatial utilization of the CGRA, reducing runtime by up to 6.2×.
Keywords:
Tensors
Pipeline processing
Micromechanical devices
Repeaters
Random access memory
Arrays
Sparse matrices
Process control
Hardware acceleration
Product design

Journal

IEEE Micro cover
IEEE Micro
IF:
2.9
Papers:
113
Citations:
2.7K

Organization

S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W