arrow
返回

SAVE: Encoding spatial interactions for vision transformers

delete2024-12-01
delete0
PRE
AI
X
Xiao Ma
张
张泽天 (Zetian Zhang)
R
Rong Yu
纪
纪则轩 (Zexuan Ji)
M
Mingchao Li
Y
Yuhan Zhang
Q
Qiang Chen *
DOI:10.1016/j.imavis.2024.105312delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Transformers have achieved impressive performance in visual tasks. Position encoding, which equips vectors (elements of input tokens, queries, keys, or values) with sequence specificity, effectively alleviates the lack of permutation relation in transformers. In this work, we first clarify that both position encoding and additional position-specific operations will introduce positional information when participating in self-attention. On this basis, most existing position encoding methods are equivalent to special affine transformations. However, this encoding method lacks the correlation of vector content interaction. We further propose Spatial Aggregation Vector Encoding (SAVE) that employs transition matrices to recombine vectors. We design two simple yet effective modes to merge other vectors, with each one serving as an anchor. The aggregated vectors control spatial contextual connections by establishing two-dimensional relationships. Our SAVE can be plug- and-play in vision transformers, even with other position encoding methods. Comparative results on three image classification datasets show that the proposed SAVE performs comparably to current position encoding methods. Experiments on detection tasks show that the SAVE improves the downstream performance of transformer-based methods. Code is available at https://github.com/maxiao0234/SAVE.
Keyword:
Vision transformers
Position encoding
Spatial interactions

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

C
china electronics technology group
学者数:
1.8K
论文数: 1.4K
被引数: 0
C
Chinese University of Hong Kong
学者数:
3.4W
论文数: 3.2W
被引数: 5.6W
引用论文

引用论文

Elimination Strategy for Aromatic Acetylenes
err2006-08-26
err0
PREAI
errAkihiro Orita; Junzo Otera
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容