arrow
Return

SPG-Mask: Fine-Grained visual classification with Self-Supervised foreground structural prior driven dynamic mask optimization

delete2026-03-03
delete0
PRE
AI
H
Huiming He
杨静 cover
杨静 (H. J. Yang)
Y
Yao Li
F
Fuyuan Cao
DOI:10.1016/j.knosys.2026.115678delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Vision Transformer (ViT) has established itself as a dominant backbone for Fine-Grained Visual Classification (FGVC). However, it encounters an intrinsic limitation wherein the global receptive field introduces background interference while the self-attention mechanism disperses focus on discriminative regions, leading to local feature sparsity. Conventional approaches typically employ hard selection strategies to filter redundancy; nevertheless, such methods often compromise the structural integrity of objects or depend upon static external cues. To overcome these limitations, a novel Foreground Structure Prior-Guided Dynamic Mask Optimization Network (SPG-Mask) is proposed. By transitioning from passive selection to active optimization, the network first constructs a latent Adaptive Foreground Structure Prior (AFSP) via self-supervised learning to capture holistic structural information. Guided by this prior, a Personalized Mask Extraction (PME) module is introduced which incorporates a differentiated reward and punishment mechanism. This mechanism dynamically accentuates structural components while suppressing background noise for individual samples, thereby ensuring the retention of minimal yet semantically integral discriminative tokens. Furthermore, a Multi-Level Token Synergy (MLTS) module is devised to integrate complementary semantics across layers. Extensive experiments conducted on three benchmark datasets, namely CUB-200-2011, Stanford Dogs, and NABirds, demonstrate that the proposed method achieves SOTA performance. The results indicate that SPG-Mask surpasses recent methods in terms of both interpretability and efficiency. The source code is available at https://github.com/Hehuiming123/SPG-Mask .
Keywords:
Vision Transformer
Fine-Grained Visual Classification
Dynamic Mask Optimization
Foreground Structure Prior
Self-Supervised Learning

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

C
computer and information technology
Scholars:
41
Papers: 18
Citations: 0