arrow
Return

LearnMat: Semantic-Aware Self-Supervision Fine-Grained Visual Recognition

delete2026-03-12
delete0
PRE
AI
S
Shuaiheng Li
F
Fan Zhang
Y
Yangyang Shu
G
Guanbin Li
J
Junyu Dong
L
Lingqiao Liu
章典 cover
章典 (David Zhang)
DOI:10.1109/TIP.2026.3671661delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Self-supervised learning has shown potential in fine-grained visual recognition (FGVR). However, existing self-supervised learning methods are often susceptible to irrelevant patterns during training and lack the ability to capture the critical subtle differences in FGVR, leading to suboptimal performance. Moreover, existing approaches focus primarily on uni-modal visual concepts. Despite the emergence of powerful vision-language models (VLMs) in various high-level vision tasks, their potential in self-supervised FGVR remains largely unexplored. To this end, we propose a novel self-supervised learning (LearnMat) framework, that effectively filters out irrelevant feature interference and extracts more important and subtle discriminative features during training. Specifically, LearnMat consists of two key modules: the semantic awareness module (SAM) and the insight extraction module (IEM). In the SAM, we introduce a novel vision–language–grounded semantic distillation strategy using a corpus of generic, category-agnostic textual attributes, that injects explicit semantic constraints into self-supervised training and improves robustness to background interference. Complementarily, the IEM exploits gradient-based signals from the input image to highlight subtle differences and localize key discriminative regions, mitigating inter-class similarity and intra-class variation, and enhancing fine-grained discrimination. Extensive experiments across multiple popular FGVR datasets show that LearnMat significantly outperforms recent state-of-the-art methods, highlighting its marked effectiveness. Our code is avaliable at https://github.com/Heng-CHY/LearnMat
Keywords:
Self-supervised learning
fine-grained visual recognition (FGVR)
discriminative feature extraction

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

O
Ocean University of China
Scholars:
3.3K
Papers: 1.1K
Citations: 2.6W
H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35
T
the chinese university of hong kong
Scholars:
4.2K
Papers: 2.0K
Citations: 0
T
the hong kong polytechnic university
Scholars:
5.0K
Papers: 2.7K
Citations: 0
S
sun yat-sen university
Scholars:
1.9W
Papers: 6.4K
Citations: 14
U
university of adelaide
Scholars:
1.2K
Papers: 757
Citations: 0
U
university of new south wales
Scholars:
2.7K
Papers: 1.4K
Citations: 0
researcher View more organizations