Return
SPG-Mask: Fine-Grained visual classification with Self-Supervised foreground structural prior driven dynamic mask optimization
DOI:10.1016/j.knosys.2026.115678.png)
Abstract
En 中文
The Vision Transformer (ViT) has established itself as a dominant backbone for Fine-Grained Visual Classification (FGVC). However, it encounters an intrinsic limitation wherein the global receptive field introduces background interference while the self-attention mechanism disperses focus on discriminative regions, leading to local feature sparsity. Conventional approaches typically employ hard selection strategies to filter redundancy; nevertheless, such methods often compromise the structural integrity of objects or depend upon static external cues. To overcome these limitations, a novel Foreground Structure Prior-Guided Dynamic Mask Optimization Network (SPG-Mask) is proposed. By transitioning from passive selection to active optimization, the network first constructs a latent Adaptive Foreground Structure Prior (AFSP) via self-supervised learning to capture holistic structural information. Guided by this prior, a Personalized Mask Extraction (PME) module is introduced which incorporates a differentiated reward and punishment mechanism. This mechanism dynamically accentuates structural components while suppressing background noise for individual samples, thereby ensuring the retention of minimal yet semantically integral discriminative tokens. Furthermore, a Multi-Level Token Synergy (MLTS) module is devised to integrate complementary semantics across layers. Extensive experiments conducted on three benchmark datasets, namely CUB-200-2011, Stanford Dogs, and NABirds, demonstrate that the proposed method achieves SOTA performance. The results indicate that SPG-Mask surpasses recent methods in terms of both interpretability and efficiency. The source code is available at https://github.com/Hehuiming123/SPG-Mask .
Keywords:
Vision Transformer
Fine-Grained Visual Classification
Dynamic Mask Optimization
Foreground Structure Prior
Self-Supervised Learning
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

