1
Return

NeuroUNI: A Unified Event-Driven Multi-Core Architecture Optimizing Neuromorphic Primitives for Brain-Inspired Computing

delete2026-05-13
delete0
PRE
AI
F
Faquan Chen
Q
Qingyang Tian
Z
Ziren Wu
X
Xiangcheng Shi
R
Rendong Ying
F
Fei Wen
P
Peilin Liu
DOI:10.1109/tcsi.2026.3690710delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The hardware convergence of Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) is hindered by conflicting computational paradigms: dense tensor parallelism versus asynchronous sparse dynamics. Existing unifications typically rely on inefficient spatial partitioning or mode-reconfigurable datapaths, limiting the flexibility needed by heterogeneous ANN-SNN hybrid models requiring frequent cross-domain interaction. To resolve this, we present NeuroUNI, a unified event-driven multi-core architecture. Unlike partitioned designs, NeuroUNI unifies computation at the primitive level using a novel Five-Tuple Event Model, abstracting both continuous activations and discrete spikes to naturally leverage dynamic sparsity. The architecture features a co-optimized hierarchical Macro-Micro-<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\mu $ </tex-math></inline-formula>OP ISA, a superscalar SIMD-based microarchitecture, and a decentralized multi-core synchronization protocol. Validated in TSMC 28nm technology via post-synthesis simulation and on a Xilinx VCU129 FPGA prototype, NeuroUNI demonstrates competitive cross-paradigm efficiency. It achieves <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$35.0\times $ </tex-math></inline-formula> and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.21\times $ </tex-math></inline-formula> the ANN energy efficiency of the NVIDIA V100 and EyerissV2, respectively, while delivering <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$5.7\times $ </tex-math></inline-formula> the SNN throughput of TrueNorth. In a unified mapless navigation workload, NeuroUNI attains 422.6 GOPS/W (ANN) and 190.8 GSOPS/W (SNN), outperforming Loihi 1 with <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$56.7\times $ </tex-math></inline-formula> the throughput and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$3.85\times $ </tex-math></inline-formula> the energy efficiency, proving the viability of a primitive/ISA-level unified silicon substrate.
Keywords:
Event-driven
ANNs
SNNs
multi-core architecture
decentralized
neuromorphic computing

Journal

IEEE Transactions on Circuits and Systems I-Regular Papers cover
IEEE Transactions on Circuits and Systems I-Regular Papers
IF:
5.2
Papers:
9.7K
Citations:
2.2W

Organization

S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
Cited Papers

Cited Papers

Citing Papers

Citing Papers