arrow
返回

TPQA: Efficient attention architecture with task-aware pattern-guided quantization

delete2026-01-21
delete0
PRE
AI
S
Sijia Wang
S
Shengbing ZHANG
L
Lun Zhang
Y
Yichao Yuan
Y
Yawen Zhao
X
Xinyu Zhang
M
Meng Zhang *
DOI:10.1016/j.future.2025.108352delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Attention mechanisms have become a cornerstone of modern deep learning models, yet their computational intensity poses significant deployment challenges for resource-limited devices. While quantization offers a potential solution, current approaches typically employ uniform precision assignment schemes across all attention heads, neglecting critical variations in head-specific contributions across different tasks. This oversight results in substantial computational redundancy for those attention heads with fewer contributions, impacting overall performance. Through systematic analysis of head pattern characteristics in transformer models, we reveal two key insights: different attention heads exhibit distinct task-aware patterns, and their varying contributions to model performance directly dictate differentiated quantization demands across heads. Building on these findings, we propose TPQA, a novel algorithm and accelerator co-design architecture for efficient deployment of transformer models. TPQA strategically assigns adaptive precision levels to each head based on pre-identified patterns, thereby reducing computational overhead while preserving model accuracy. Furthermore, TPQA employs a data reordering strategy to transform irregular workloads into structured formats and introduces a dedicated accelerator with an attention-weights-stationary dataflow to efficiently process these structured workloads. Comprehensive evaluations demonstrate TPQA's superior performance, achieving up to 2.1x speedup and 3.4x energy efficiency improvement over state-of-the-art accelerators while maintaining <1% accuracy loss on various tasks.
Keyword:
Multi-head attention
Quantization
Accelerator
Head pattern
Systolic array
Attention architecture

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

N
northwestern polytechnical university
学者数:
1.3W
论文数: 4.6K
被引数: 0
引用论文

引用论文

Mixed-Precision Post-Training Quantization for Learned Image Compression
err2025-08-15
err0
PREAI
errYu,Jie; Mai,Songping; Zhang,Peng; Jiang,Yucheng; Cheng,Jian
err分享
err收藏
Hardware–Software Co-Design Enabling Static and Dynamic Sparse Attention Mechanisms
err
IF0
err2024-09-01
err0
PREAI
errJieru Zhao; Pai Zeng; Guan Shen; Quan Chen; Minyi Guo
err分享
err收藏
Ramulator 2.0: A modern, modular, and extensible dram simulatorRamulator 2.0:一款现代、模块化且可扩展的DRAM模拟器
err
IF0
err2024-01-01
err0
PREAI
errHaocong Luo; Yahya Can Tuğrul; F. Nisa Bostancı; Ataberk Olgun; A. Giray Yağlıkçı; Onur Mutlu
err分享
err收藏
Block-Wise Mixed-Precision Quantization: Enabling High Efficiency for Practical ReRAM-Based DNN Accelerators分块混合精度量化:实现实用ReRAM基DNN加速器的高效性
err2024-12-01
err0
errOAAI
errXueying Wu; Edward Hanson; Nansu Wang; Qilin Zheng; Xiaoxuan Yang; Huanrui Yang; Shiyu Li; Feng Cheng; Partha Pratim Pande; Janardhan Rao Doppa; Krishnendu Chakrabarty; Hai Li
err分享
err收藏
学者 查看更多内容