Return
NeuroUNI: A Unified Event-Driven Multi-Core Architecture Optimizing Neuromorphic Primitives for Brain-Inspired Computing
F
Q
Z
X
R
F
P
DOI:10.1109/tcsi.2026.3690710.png)
Abstract
En 中文
The hardware convergence of Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) is hindered by conflicting computational paradigms: dense tensor parallelism versus asynchronous sparse dynamics. Existing unifications typically rely on inefficient spatial partitioning or mode-reconfigurable datapaths, limiting the flexibility needed by heterogeneous ANN-SNN hybrid models requiring frequent cross-domain interaction. To resolve this, we present NeuroUNI, a unified event-driven multi-core architecture. Unlike partitioned designs, NeuroUNI unifies computation at the primitive level using a novel Five-Tuple Event Model, abstracting both continuous activations and discrete spikes to naturally leverage dynamic sparsity. The architecture features a co-optimized hierarchical Macro-Micro-<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\mu $ </tex-math></inline-formula>OP ISA, a superscalar SIMD-based microarchitecture, and a decentralized multi-core synchronization protocol. Validated in TSMC 28nm technology via post-synthesis simulation and on a Xilinx VCU129 FPGA prototype, NeuroUNI demonstrates competitive cross-paradigm efficiency. It achieves <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$35.0\times $ </tex-math></inline-formula> and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.21\times $ </tex-math></inline-formula> the ANN energy efficiency of the NVIDIA V100 and EyerissV2, respectively, while delivering <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$5.7\times $ </tex-math></inline-formula> the SNN throughput of TrueNorth. In a unified mapless navigation workload, NeuroUNI attains 422.6 GOPS/W (ANN) and 190.8 GSOPS/W (SNN), outperforming Loihi 1 with <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$56.7\times $ </tex-math></inline-formula> the throughput and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$3.85\times $ </tex-math></inline-formula> the energy efficiency, proving the viability of a primitive/ISA-level unified silicon substrate.
Keywords:
Event-driven
ANNs
SNNs
multi-core architecture
decentralized
neuromorphic computing
Journal
IF:
5.2
Papers:
9.7K
Citations:
2.2W
