Return
Clustering in Pure-Attention Hardmax Transformers and Its Role in Sentiment Analysis
A
G
E
DOI:10.1137/24M167086X.png)
Abstract
En 中文
Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers with hardmax self-attention and normalization sublayers as the number of layers tends to infinity. By viewing such transformers as discrete-time dynamical systems describing the evolution of points in a Euclidean space, and thanks to a geometric interpretation of the self-attention mechanism based on hyperplane separation, we show that the transformer inputs asymptotically converge to a clustered equilibrium determined by special points called leaders. We then leverage this theoretical understanding to solve sentiment analysis problems from language processing using a fully interpretable transformer model, which effectively captures ``context by clustering meaningless words around leader words carrying the most meaning. Finally, we outline remaining challenges to bridge the gap between the mathematical analysis of transformers and their real-life implementation.
Keywords:
transformers
self-attention
clustering
AI interpretability
sentiment analysis
Journal
S
IF:
2.6
Papers:
17
Citations:
0
