Return
KV-Cache Oriented Query-Aware Sparse Attention Accelerator With Cross-Stage Precision-Configurable Digital CIM
Y
W
Y
L
张
R
H
B
DOI:10.1109/TCSII.2025.3580135.png)
Abstract
En 中文
This brief proposes KV-CIM, a KV-Cache oriented Digital Compute-In-Memory (DCIM) sparse attention accelerator, to address computational and memory bottlenecks in autoregressive inference for large language models. Key innovations include: a) A query-aware pre-compute architecture, which dynamically selects and accesses KV-Cache for critical tokens at the pre-compute stage (Stage1) and deploys KV-Cache segmentally on memory-constrained edge devices while maintaining computational accuracy at the formal computation stage (Stage2); b) A cross-stage DCIM macro featuring precision-configurable adder trees, which works in approximate mode at Stage1 and changes to full precision mode at Stage2; c) A query-stationary dataflow that retains the current query tensors in q-CIM across stages to eliminate data movement. Under 28-nm CMOS technology, the proposed KV-CIM achieves 35.16 TOPS/W and 82% reduction of external memory access with negligible degradation in LLaMA2 expressiveness.
Keywords:
Compute-in-memory based accelerator
KV-cache
sparse attention
precision-configurable DCIM
Journal
I
IF:
4.9
Papers:
8.8K
Citations:
2.5W
