arrow
Return

Spatial-frequency cross-shift learning perceptive transformer features in decoupled hashing

delete2025-08-18
delete0
PRE
AI
J
Jing Zhang
S
Shuli Cheng
L
Liejun Wang *
DOI:10.1016/j.patcog.2025.112287delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the rapid development of multimedia technology, the number of images has grown significantly, making the search for similar images an urgent necessity in everyday life. Hash image retrieval has gradually dominated the field of image retrieval due to its advantages of computational efficiency and high accuracy. However, existing image retrieval algorithms, whether based on convolutional neural networks (CNN) or Vision Transformers (ViT), primarily focus on enhancing the extraction of target category features in the spatial domain, neglecting the role of features in other domains, which to some extent limits retrieval accuracy. This paper proposes a Spatial-Frequency Cross-Shift Learning Perceptive Transformer Features in Decoupled Hashing (SFPTH), for image retrieval. First, based on the spatial location distribution of image features and the target category features present in the channels, we designed the Spatial-Spectral Dual-Domain Shift Perception (SSD) module within deep semantic feature layers. This design not only focus on features of the same category across different locations, but also the embedded frequency domain can also learn category information brought by target features enhanced through spatial shifts, thereby strengthening the output in deep features. This provides rich semantic information for discrete mapping, reducing the probability of misclassification. Secondly, we designed a bit decoupling loss to promote the independence of bits, which reduces the similarity of discrete values in model output by penalizing the correlation between labels, ensuring that the binarized output of the hash codes is more discriminative. We conducted extensive experiments on three public image retrieval datasets: MS-COCO, ImageNet, and CIFAR10, achieving retrieval performance of 93.33%, 94.92%, and 94.87%, respectively.
Keywords:
image retrieval
hash coding
deep learning
spatial-frequency domain
feature decoupling

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

No organization information available