1
Return

PointTFA<inline-formula><tex-math notation="LaTeX">$^{m}$</tex-math></inline-formula>: Multi-Modal, Training-Free Adaptation for Point Cloud Understanding

delete2026-03-16
delete0
PRE
AI
J
Jinmeng Wu
Y
Youxiang Hu
C
Chong Cao
H
Hao Zhang
B
Basura Fernando
Y
Yanbin Hao
H
Hanyu Hong
DOI:10.1109/tmm.2026.3668609delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High-dimensional data are more sparsely distributed in space compared to low-dimensional data of the same size (e.g., 3D point cloud <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vs</i> 2D images), a phenomenon known as the “<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Curse of Dimensionality</i>” (COD). Consequently, more samples are required to effectively fine-tune models for high-dimensional tasks like 3D point cloud understanding, leading to increased computational costs. Meanwhile, although 3D point clouds provide comprehensive spatial details, 2D images projected from specific viewpoints often capture sufficient information for understanding visual content. To address the COD challenge and leverage the complementary nature of 3D-2D data, we introduce a multi-modal, training-free approach named PointTFA<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{m}$</tex-math></inline-formula>, an extended version of our original PointTFA. This new approach incorporates 2D view images projected from 3D point clouds in training-free manner to augment cloud classification. Specifically, PointTFA<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{m}$</tex-math></inline-formula> contains two training-free branches that process 3D point clouds and 2D view images independently. Each branch includes its own Representative Memory Cache (RMC), Cloud/Image Query Refactor (CQR or IQR), and Training-Free Adapter (TFA). The model combines the outputs from both branches through score fusion to make effective multi-modal predictions. PointTFA<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{m}$</tex-math></inline-formula> improves upon single-modal PointTFA by accuracy gains of 1.01%, 1.32%, and 4.64% on the ModelNet40, ModelNet10, and ScanObjectNN benchmarks, respectively, setting new state-of-the-art performance for training-free point cloud understanding approaches.
Keywords:
Multimodal fusion
3D visual understanding
few-shot learning
training-free adaption

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

A
agency for science, technology research
Scholars:
2
Papers: 1
Citations: 0
H
Hefei University of Technology
Scholars:
4.7K
Papers: 1.6K
Citations: 2.1W
A
agency for science, technology and research
Scholars:
497
Papers: 194
Citations: 0
W
wuhan institute of technology
Scholars:
9.8K
Papers: 6.4K
Citations: 11
Cited Papers

Cited Papers

Citing Papers

Citing Papers