arrow
Return

S3ViT: Self-Supervised Spectral Vision Transformer Framework for Hyperspectral Unmixing

delete2026-04-22
delete0
PRE
AI
D
DS Dario Scilla
V
VA Victor Angulo
K
KJ Kasper Johansen
N
NA Naif Alsalem
W
Wolfgang Heidrich
M
Matthew F. McCabe
DOI:10.3389/frsen.2026.1812755delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Hyperspectral unmixing aims to decompose each pixel in a hyperspectral image into a set of constituent endmembers and their corresponding abundances. Recent deep learning based approaches have demonstrated strong performance in capturing both spectral and spatial features. However; obtaining reliable per-pixel abundance ground truth in real hyperspectral scenes is generally infeasible; which motivates unsupervised and self-supervised unmixing strategies. In this work; we propose S3ViT; a self-supervised Spectral Vision Transformer designed for pixel-wise hyperspectral unmixing. The transformer captures spectral and spatial dependencies by applying self-attention over the full sequence of pixel tokens (1×1) augmented with learnable positional embeddings. It operates without ground-truth annotations by generating pseudo labels through an unsupervised process: first; Singular Value Decomposition (SVD) is used to estimate the number of endmembers based on a thresholded singular value spectrum; then; k-means clustering provides cluster-derived priors that are used to form a contextual token and initialize spectral prototypes; without being treated as true abundance supervision. To guide training; two initialization tokens; one from Vertex Component Analysis (VCA) and one from the k-means cluster-derived priors; are embedded alongside patch tokens into the transformer. The model learns to estimate abundance maps while enforcing the Abundance Non-negativity and Sum-to-One constraints through ReLU and Softmax layers. Endmember spectra are later estimated from pixels with high predicted abundances. We evaluated S3ViT on the Samson; Jasper Ridge and Washington DC Mall benchmark datasets and compared it to state-of-the-art geometrical and deep learning methods. Our model achieves superior or comparable performance in both RMSE and SAD metrics; with up to 31% improvement in SAD and 25% in RMSE. These results indicate that a compact pixel-token ViT; guided by weak spectral priors and optimized via reconstruction losses; can achieve competitive unmixing performance on standard benchmarks.
Keywords:
Hyperspectral imagery
Hyperspectral unmixing
remote sensing
Self-supervised
Visual Transformer (ViT)

Journal

F
Frontiers in Remote Sensing
IF:
3.7
Papers:
569
Citations:
993

Organization

K
king abdullah university of science and technology
Scholars:
1.6K
Papers: 563
Citations: 0