Return
Concatenation-based positional encoding for transformer models: proposal and performance analysis
DOI:10.5351/KJAS.2026.39.2.179.png)
Abstract
En 中文
Despite their success across domains, Transformer models face challenges in time-series forecasting due to their permutation-invariant attention mechanism, which neglects positional dependencies. Traditional positional encoding alleviates this issue, but its element-wise addition to input embeddings often causes interference between semantic and positional information. This study investigates the structural and theoretical characteristics of concatenation-based positional encoding in comparison with the conventional addition-based approach. By analyzing the attention score formulation, we show that concatenation-based encoding enables semantic and positional components to be processed in independent subspaces, thereby introducing a different inductive bias in the attention mechanism. Attention map visualizations are further employed to qualitatively examine how positional information is reflected under each encoding structure, providing insights into their interpretability. Empirical evaluations are conducted on both time-series forecasting and natural language processing tasks to examine performance and interpretability. The results show that concatenation-based positional encoding yields improved performance compared to the addition-based approach across both task domains, with more noticeable gains in scenarios where positional information plays a critical role. Through a systematic analysis of structural behavior and task-dependent effects, this work contributes to a clearer understanding of how positional information is handled in Transformer models.
Keywords:
transformer models
positional encoding
concatenation
Journal
K
IF:
0
Papers:
20
Citations:
0

