1
Return

KV-Cache Oriented Query-Aware Sparse Attention Accelerator With Cross-Stage Precision-Configurable Digital CIM

delete2025-08-01
delete0
PRE
AI
Y
Yang Zhang
W
Weixuan Wang
Y
Yizhi Ding
L
Lizheng Ren
张毅然 (Yiran Zhang)
R
Ruiqi Tan
王振 cover
王振 (Zhen Wang)
H
Hao Cai
B
Bo Liu
DOI:10.1109/TCSII.2025.3580135delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This brief proposes KV-CIM, a KV-Cache oriented Digital Compute-In-Memory (DCIM) sparse attention accelerator, to address computational and memory bottlenecks in autoregressive inference for large language models. Key innovations include: a) A query-aware pre-compute architecture, which dynamically selects and accesses KV-Cache for critical tokens at the pre-compute stage (Stage1) and deploys KV-Cache segmentally on memory-constrained edge devices while maintaining computational accuracy at the formal computation stage (Stage2); b) A cross-stage DCIM macro featuring precision-configurable adder trees, which works in approximate mode at Stage1 and changes to full precision mode at Stage2; c) A query-stationary dataflow that retains the current query tensors in q-CIM across stages to eliminate data movement. Under 28-nm CMOS technology, the proposed KV-CIM achieves 35.16 TOPS/W and 82% reduction of external memory access with negligible degradation in LLaMA2 expressiveness.
Keywords:
Compute-in-memory based accelerator
KV-cache
sparse attention
precision-configurable DCIM

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

S
Southeast University
Scholars:
1.8W
Papers: 7.7K
Citations: 480
Cited Papers

Cited Papers

Citing Papers

Citing Papers