1
Return

MDS-ViT: A multi-data stream FPGA-Based vision transformer accelerator

delete2026-06-19
delete0
PRE
AI
W
Wenhua Ye
H
Huan Li *
X
Xu Zhou
D
Dong Pan
K
Kenli Li
DOI:10.1016/j.sysarc.2026.103898delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• A Novel DDR-Centric, Multi-Data-Stream Architecture for Scalability. We propose a new design paradigm that re-architects the ViT computation around high-bandwidth DDR memory, rather than treating it as a bottleneck. The core is a deeply pipelined “split-compute-aggregate” dataflow. This architecture intrinsically decouples processing capability from on-chip memory capacity, enabling efficient acceleration of models from small (DeiT-Tiny) to extremely large (ViT-Huge) scales. The implementation of 4 parallel HP_MMs within an SLR demonstrates the practical viability of this paradigm, achieving a 4x performance gain. • A Configurable, Model-Agnostic Acceleration Flow. To achieve the claimed versatility, we introduce a programmable acceleration process controlled via HLS DATAFLOW and configuration files. This abstraction decouples the hardware pipeline from specific model dataflows, allowing the same physical architecture to dynamically adapt to different ViT/DeiT variants by dividing the workflow into 10 reconfigurable steps. This provides a level of flexibility unattainable by static, fixed-function accelerators. • A Unified Compute Core with a Dual-Interface DDR Strategy. We design a high-performance matrix multiplication (HP_MM) core that is agile to heterogeneous data sources, coupled with a novel dual-interface DDR reading method. This co-design allows the computational core to natively support the diverse matrix operations in MHA and MLP (e.g., implicitly handling transpositions by swapping data interfaces), eliminating the need for dedicated, resource-intensive transpose modules and streamlining the data supply for critical computations. • An Intelligent Multi-Level Cache Hierarchy as a Bandwidth Amplifier. Recognizing that DDR bandwidth is the ultimate system bottleneck in a DDR-centric design, we architect the on-chip cache not merely as storage, but as an active bandwidth amplification and access pattern conversion layer. This hierarchy, particularly the two-level ping-pong buffering for weights, effectively mitigates DDR access latency, transforming bursty, off-chip memory traffic into a steady, high-throughput stream for the computing array, which is the key to making the DDR-centric approach efficient. • Experimental results demonstrate that under the ViT-Base model, MDS-ViT achieves energy consumption reductions of 27.91 × and 2.82 × compared to CPU and GPU, respectively. Under the ViT-Huge model, these savings are 26.99 × and 3.11 × relative to CPU and GPU, respectively.

Journal

Journal of Systems Architecture cover
Journal of Systems Architecture
IF:
4.1
Papers:
2.9K
Citations:
4.2K

Organization

H
hunan university
Scholars:
4.3W
Papers: 3.2W
Citations: 70
Cited Papers

Cited Papers

Citing Papers

Citing Papers