Return
PAT-Net: Point agent transformer network for point cloud classification
DOI:10.1016/j.knosys.2026.115851.png)
Abstract
En 中文
As the cornerstone of 3D vision, point cloud classification remains a challenging task. To tackle the quadratic cost of self-attention in the number of points, most existing Transformer-based methods compute self-attention only within individual patches or clusters of the point cloud, but this inevitably limits the ability to model global contexts. Furthermore, these methods fail to effectively learn both high- and low-frequency features together from the input point clouds. To overcome the above issues, in this paper, we present Point Agent Transformer Network (PAT-Net), a novel network for point cloud classification. Specifically, we first propose an offset-agent Transformer (OAT) block, which employs a small set of agent points that serve as intermediaries to aggregate and broadcast global information, achieving reduced complexity while preserving global context modeling capability. Besides, we design a dual-frequency neighborhood feature aggregation (DNFA) module and a feature enhancement fusion (FEF) module, which work in tandem to effectively learn comprehensive features with both high- and low-frequency information in point clouds. With the DNFA module and the FEF module as fundamental components, we construct the PAT-Net for 3D point cloud classification. Extensive experiments demonstrate that our method achieves state-of-the-art performance on ScanObjectNN, ModelNet40 and ModelNet-O datasets.
Keywords:
Point cloud classification
Transformer
Global context modeling
Dual-frequency features
Feature enhancement
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
Organization
No organization information available

