1
Return

SalFormer360: A Transformer-Based Saliency Estimation Model for 360-Degree Videos

delete2026-03-06
delete0
delete
OA
AI
M
Mahmoud Z. A. Wahba
F
Francesco Barbato
S
Sara Baldoni
F
Federica Battisti
DOI:10.1109/tbc.2026.3668621delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Saliency estimation has received growing attention in recent years due to its importance in a wide range of applications. In the context of 360-degree video, it has been particularly valuable for tasks such as viewport prediction and immersive content optimization. In this paper, we propose SalFormer360, a novel saliency estimation model for 360-degree videos built on a transformer-based architecture. Our approach is based on the combination of an existing encoder architecture, SegFormer, and a custom decoder. The SegFormer model was originally developed for 2D segmentation tasks, and it has been fine-tuned to adapt it to 360-degree content. To further enhance prediction accuracy in our model, we incorporated a viewing center bias to reflect user attention in 360-degree environments. Extensive experiments on the three largest benchmark datasets for saliency estimation demonstrate that SalFormer360 outperforms existing state-of-the-art methods. In terms of Pearson correlation coefficient, our model achieves 8.4% higher performance on Sport360, 2.5% on PVS-HM, and 18.6% on VR-EyeTracking compared to previous state-of-the-art.
Keywords:
Saliency estimation
omni-directional video
viewing bias
transformers

Journal

IEEE Transactions on Broadcasting cover
IEEE Transactions on Broadcasting
IF:
4.8
Papers:
2.1K
Citations:
3.0K

Organization

U
university of Padova
Scholars:
1.2K
Papers: 432
Citations: 10
Cited Papers

Cited Papers

Citing Papers

Citing Papers