arrow
Return

CIMFlow: Modelling Dataflow in Cross-Layer Compute-in-Memory Deep Learning Accelerators

delete2025-09-01
delete0
delete
OA
AI
J
José Cubero-Cascante *
S
Schneider, Lucas Tonini Rosenberg
R
Rebecca Pelke
A
Arunkumar Vaidyanathan
R
Rainer Leupers
J
Jan Moritz Joseph
DOI:10.1145/3760780delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Traditional Deep Learning Accelerators (DLAs) rely on off-chip memory to store large weight tensors, leading to high bandwidth demands and energy consumption. Compute-in-Memory (CIM) accelerators mitigate this by integrating high-density, non-volatile memory arrays, enabling a fully weight-stationary (FWS) dataflow. Multi-core CIM systems further enhance efficiency with cross-layer inference, where intermediate tensors stay on-chip, and cores operate in a pipeline. Despite diverse architecture proposals, no existing tool models the dataflow, memory access patterns and timing behaviour of multi-core CIM accelerators. We introduce CIMFlow, a modelling framework for cross-layer CIM architectures. Our flexible Hardware Architecture Model includes CIM and digital cores and leverages the buffets storage idiom for distributed token-based flow control. An Array-OL-based Workload Model captures CNNs' multidimensional dependencies and applies hardware-aware transformations. These models are transformed into a timed cyclo-static dataflow graph for simulation. CIMFlow delivers latency, energy and traces for core and buffer utilisation. Our case studies on state-of-the-art CNNs show that cross-layer inference reduces latency by up to 52x. We also reveal that neglecting memory access delays results in throughput overestimations of up to 308%. To our knowledge, CIMFlow is the first tool focused on FWS cross-layer execution in CIM architectures that explicitly models data movement costs. It serves as a powerful cost model for design space exploration in next-generation CIM accelerators.
Keywords:
Compute-in-memory
cross-layer
dataflow simulation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Embedded Computing Systems cover
ACM Transactions on Embedded Computing Systems
IF:
2.6
Papers:
227
Citations:
2.3K

Organization

R
RWTH Aachen University
Scholars:
3.5W
Papers: 2.6W
Citations: 3.6W
Cited Papers

Cited Papers

A Calibratable Model for Fast Energy Estimation of MVM Operations on RRAM Crossbars
err2024-04-22
err0
errOAAI
errJosé Cubero-Cascante; Arunkumar Vaidyanathan; Rebecca Pelke; Lorenzo Pfeifer; Rainer Leupers; Jan Moritz Joseph
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
MTIA: First Generation Silicon Targeting Meta's Recommendation Systems
err2023-06-17
err0
PREAI
errAmin Firoozshahian; Joel Coburn; Roman Levenstein; Rakesh Nattoji; Ashwin Kamath; Olivia Wu; Gurdeepak Grewal; Harish Aepala; Bhasker Jakka; Bob Dreyer; Adam Hutchin; Utku Diril; Krishnakumar Nair; Ehsan K. Aredestani; Martin Schatz; Yuchen Hao; Rakesh Komuravelli; Kunming Ho; Sameer Abu Asal; Joe Shajrawi; Kevin Quinn; Nagesh Sreedhara; Pankaj Kansal; Willie Wei; Dheepak Jayaraman; Linda Cheng; Pritam Chopda; Eric Wang; Ajay Bikumandla; Arun Karthik Sengottuvel; Krishna Thottempudi; Ashwin Narasimha; Brian Dodds; Cao Gao; Jiyuan Zhang; Mohammed Al-Sanabani; Ana Zehtabioskuie; Jordan Fix; Hangchen Yu; Richard Li; Kaustubh Gondkar; Jack Montgomery; Mike Tsai; Saritha Dwarakapuram; Sanjay Desai; Nili Avidan; Poorvaja Ramani; Karthik Narayanan; Ajit Mathews; Sethu Gopal; Maxim Naumov; Vijay Rao; Krishna Noru; Harikrishna Reddy; Prahlad Venkatapuram; Alexis Bjorlin
errShare
errSave
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
err2025-01-01
err0
PREAI
errSymons, Arne; Mei, Linyan; Colleman, Steven; Houshmand, Pouya; Karl, Sebastian; Verhelst, Marian
errShare
errSave
Buffets
err2019-04-04
err0
errOAAI
errMichael Pellauer; Yakun Sophia Shao; Jason Clemons; Neal Crago; Kartik Hegde; Rangharajan Venkatesan; Stephen W. Keckler; Christopher W. Fletcher; Joel Emer
errShare
errSave
A Heterogeneous and Programmable Compute-In-Memory Accelerator Architecture for Analog-AI Using Dense 2-D Mesh
err2023-01-01
err0
PREAI
errShubham Jain; Hsinyu Tsai; Ching-Tzu Chen; Ramachandran Muralidhar; Irem Boybat; Martin M. Frank; Stanislaw Wozniak; Milos Stanisavljevic; Praneet Adusumilli; Pritish Narayanan; Kohji Hosokawa; Masatoshi Ishii; Arvind Kumar; Vijay Narayanan; Geoffrey W. Burr
errShare
errSave
researcher View more