arrow
返回

WindowKV: Task-adaptive group-wise KV cache window selection for efficient LLM inference

delete2026-09-04
delete0
PRE
AI
Y
Youhui Zuo
S
Sibo Wei
C
Chen Zhang
Z
Zhuorui Liu
W
Wenpeng Lü
宋
宋大为 (Dawei Song) *
DOI:10.1016/j.eswa.2026.134291delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
• WindowKV为保持KV一致性选择任务感知的连续语义窗口。 • 一种组内KV索引共享策略降低了开销并提升了效率。 • WindowKV使用12%的KV实现近满性能,并在LongBench和NIAH上表现优异。
Keyword:
Large language model
KV cache compression
Inference efficiency

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

Q
qilu university of technology
学者数:
2.1K
论文数: 642
被引数: 0
B
beijing institute of technology
学者数:
5.5W
论文数: 4.0W
被引数: 63
引用论文

引用论文

暂无论文信息