Return
Global-local information sensitivity adjustment factor
DOI:10.1016/j.patrec.2025.03.031.png)
Abstract
En 中文
In recent years, the dominant trend in computer vision has shifted to large receptive field networks represented by Transformers and large kernel convolution networks. However, utilizing a fixed large receptive field structure across all stages, as done by most existing methods, can potentially hinder the efficiency of feature extraction. This is because different stages of the network may require varying receptive fields based on the specific tasks. To address this issue, we propose a learnable scaling factor that modulates between the large and small receptive field structure of the network. The scaling factor is trained together with the network, enabling adaptive adjustment of the sensitivity to global-local information in different stages and tasks. By incorporating the learnable scaling factor we are able to dynamically balance the contribution of global and local features in the output feature maps with almost no increase in FLOPs and parameters. We demonstrate the effectiveness of our approach by applying this improvement to two baseline methods, RepLKNet and VAN (via the newly proposed receptive field separation convolution). Experiments demonstrate that our proposed method can improve the performance of baseline methods on image classification, object detection, and semantic segmentation. In addition, we also introduced learnable scaling factors into small kernel convolution networks as well as Transformer-based networks and observed performance improvements, further proving the robustness of the method.
Keywords:
Locality
Vision backbone
Large kernel convolution
Journal
IF:
3.3
Papers:
7.8K
Citations:
1.6W

