arrow
Return

Efficient High-Order Spatial Interactions for Visual Perception

delete2025-08-28
delete0
PRE
AI
刘祖岩 (Zuyan Liu)
Y
Yongming Rao
W
Wenliang Zhao
周杰 (Jie Zhou)
J
Jiwen Lu
DOI:10.1109/TPAMI.2025.3603181delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent progress in vision Transformers exhibits great success in various tasks driven by the new spatial modeling mechanism based on dot-product self-attention. In this paper, we show that the key ingredients behind the vision Transformers, namely input-adaptive, long-range and high-order spatial interactions, can also be efficiently implemented with a convolution-based framework. We present the Recursive Gated Convolution (<inline-formula><tex-math notation="LaTeX">${\mathit{g}}^{\mathit{n}}$</tex-math></inline-formula>Conv) that performs high-order spatial interactions with gated convolutions and recursive designs. The new operation is highly flexible and customizable, which is compatible with various variants of convolution and extends the two-order interactions in self-attention to arbitrary orders without introducing significant extra computation. <inline-formula><tex-math notation="LaTeX">${\mathit{g}}^{\mathit{n}}$</tex-math></inline-formula> Conv can serve as a plug-and-play module to improve various vision Transformers and convolution-based models. Based on the proposed operation, we construct a new family of generic vision backbones for various visual modalities and tasks, including HorNet and HorFPN for image recognition, Hor3D for point cloud analysis, and HorCLIP for vision-language modeling. For image recognition, we propose HorNet as a stronger visual encoder, where we conduct extensive experiments on ImageNet classification, COCO object detection, and ADE20K semantic segmentation. HorNet outperforms Swin Transformers and ConvNeXt by a significant margin with similar overall architecture and training configurations. HorNet also shows favorable scalability to more training data and larger model sizes. Apart from image encoders, we also show <inline-formula><tex-math notation="LaTeX">${\mathit{g}}^{\mathit{n}}$</tex-math></inline-formula>Conv can be applied to task-specific decoders and consistently improve dense prediction performance with less computation. For point cloud analysis, we design Hor3D, demonstrating the efficacy of high-order interactions for unstructured point cloud data through experiments on challenging 3D semantic segmentation tasks in S3DIS and ScanNet V2. In vision-language modeling, our proposed HorCLIP surpasses mainstream Vision Transformer and ConvNeXt architectures with shorter training schedules on ImageNet zero-shot classification and shows remarkably higher performance on vision-language dense representation tasks on COCO Panoptic datasets. Our results demonstrate that <inline-formula><tex-math notation="LaTeX">${\mathit{g}}^{\mathit{n}}$</tex-math></inline-formula>Conv with high-order spatial interactions can be a new basic operation for visual modeling that effectively combines the merits of both vision Transformers and CNNs.
Keywords:
Visual perception
high-order spatial interactions
recursive gated convolution

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 10.0W
Citations: 137