Return
Query attention: An effective plug-and-play module for enhanced small dataset classification
C
R
H
DOI:10.1016/j.neucom.2026.134731.png)
Abstract
En 中文
In small-dataset image classification, limited training samples often hinder models from learning sufficiently discriminative feature representations. Although various plug-and-play modules have been developed to enhance feature learning, relatively limited attention has been paid to explicitly refining global channel semantics and leveraging them to guide spatial feature aggregation, especially in a unified manner applicable to both convolutional neural networks (CNNs) and Vision Transformers (ViTs). To address this issue, we propose Query Attention (QA), a lightweight plug-and-play module that can be integrated into both ViTs and CNNs with minimal changes to the backbone pipelines. QA consists of two components: Query-Driven Channel Attention (QDCA) and Channel-Spatial Interaction Attention (CSIA). QDCA refines global channel semantics through the interaction between a learnable query prior and a backbone-derived global channel descriptor, while CSIA uses the refined channel representation to guide spatial feature aggregation. In addition, we introduce a pruning strategy for QA-enhanced ViTs based on attention entropy and gradient norm, yielding a compact structure that reduces computational cost while maintaining competitive performance. We integrate QA into various backbone networks, including ViT, Swin, T2T-ViT, ResNet18, MobileNetV2, and VGG19, and evaluate it on four small datasets: CIFAR-10, CIFAR-100, Tiny-ImageNet, and CINIC-10. Experimental results show that QA improves classification accuracy across different backbone families and datasets, with gains of up to 7.42%, demonstrating the effectiveness and generality of the proposed method. Code is available at: https://github.com/jcy619/QA .
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
