arrow
Return

HyperPoint: Multimodal 3D foundation model in hyperbolic space

delete2025-12-04
delete0
PRE
AI
Y
Yiding Sun
H
Haozhe Cheng
C
Chaoyi Lu
Z
Zhengqiao Li
M
Minghong Wu
路慧敏 (Huimin Lu)
祝继华 cover
祝继华 (Jihua Zhu)
DOI:10.1016/j.patcog.2025.112800delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose HyperPoint, the first multimodal 3D foundation model in hyperbolic space. Our method not only unifies generative and contrastive learning but also extends cross-modal contrastive learning to hyperbolic space. • We propose a comprehensive multi-loss optimization method to tackle the challenges of hierarchical and multi-modal complexities in point cloud representation learning. This method integrates generative, contrastive, and structural learning objectives into a unified design, ensuring a balance between feature diversity and semantic consistency. • We perform thorough qualitative analyses with HyperPoint to demonstrate its potential for capturing hierarchical structures and preserving semantic consistency in cross-modal data. Through extensive experiments, we demonstrate that the hyperbolic counterpart outperforms the Euclidean setting.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

X
xi’an jiaotong university
Scholars:
7.7K
Papers: 2.4K
Citations: 1
S
Southeast University
Scholars:
1.9W
Papers: 8.2K
Citations: 480