Return
A Runtime Reconfigurable Array (RTRA) for Multiprogram ML and DSP Acceleration With Hierarchical Partitioning and On-Chip Multilevel Scheduling
DOI:10.1109/jssc.2026.3715500.png)
Abstract
En 中文
Emerging applications with highly dynamic digital signal processing (DSP) and machine-learning (ML) workloads demand a hardware platforms that simultaneously deliver high performance, high energy efficiency, and low-latency reconfiguration. Such systems must adapt to evolving algorithms, variable data rates, and multiprogram tenancy without incurring prohibitive switching overhead. The conventional reconfigurable architectures and their associated scheduling frameworks typically rely on off-chip, software-based scheduling and mapping, introducing substantial latency and limited scalability. In addition, their dependence on the complex 2-D placement methods across processing-element (PE) arrays constrains responsiveness and resource utilization, making them unsuitable for real-time and multiprogram operation. To address these limitations, we present a runtime reconfigurable array (RTRA), a coarse-grained hardware fabric that integrates on-chip scheduling and is architected around modular, hierarchically partitioned compute blocks. This organization enables fast scheduling and mapping, sustains high utilization across diverse workloads, supports fine-grained multiprogram tenancy, and achieves high energy and area efficiency.
Keywords:
Coarse-grained reconfigurable architecture (CGRA)
network-on-chip (NoC)
on-chip hardware scheduling
runtime reconfiguration
system-on-chip (SoC)
Journal
I
IF:
5.6
Papers:
888
Citations:
2.7W

