arrow
Return

Word embedding factor based multi-head attention

delete2025-01-30
delete0
delete
OA
AI
Z
Zhengren Li
Z
Zhao, Yumeng
X
Xiaohang Zhang *
H
Huawei Han
C
Cui Fen Huang
DOI:10.1007/s10462-025-11115-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The natural language processing (NLP) field has made significant progress using deep learning models based on multi-head attention mechanisms, such as Transformer and BERT. However, there are two major limitations to this approach. First, the number of heads is often manually set based on empirical experience, and second, it is not clear enough in semantic understanding and interpretation. In this study, we propose a novel attention mechanism called Factor Analysis-based Multi-head (FAM) Attention, which combines the theory of explorative factor analysis and word embedding. The experimental results demonstrate that FAM Attention achieves better performance and requires fewer parameters compared to traditional methods while also having better semantic understanding ability and interpretability at the token level. This also has significant implications for current Large Language Models (LLMs), particularly in terms of effectively reducing parameter counts and enhancing performance.
Keywords:
Multi-head attention
Word embedding
Factor analysis
BERT

Journal

Artificial Intelligence Review cover
Artificial Intelligence Review
IF:
13.9
Papers:
6.1K
Citations:
1.9W

Organization

B
Beijing Univ Posts and Telecommun
Scholars:
719
Papers: 311
Citations: 55
Cited Papers

Cited Papers

Annotating Columns with Pre-trained Language Models
err2022-06-11
err0
errOAAI
errYoshihiko Suhara; Jinfeng Li; Yuliang Li; Dan Zhang; Çağatay Demiralp; Chen Chen; Wang-Chiew Tan
errShare
errSave
Video Swin Transformer
err2022-06-01
err0
PREAI
errZe Liu; Jia Ning; Yue Cao; Yixuan Wei; Zheng Zhang; Stephen Lin; Han Hu
errShare
errSave
On the diversity of multi-head attention
err2021-09-01
err53
PREAI
errLi, Jian; Wang, Xing; Tu, Zhaopeng; Lyu, Michael R.
errShare
errSave
researcher View more