Return
HyperPoint: Multimodal 3D foundation model in hyperbolic space
DOI:10.1016/j.patcog.2025.112800.png)
Abstract
En 中文
• We propose HyperPoint, the first multimodal 3D foundation model in hyperbolic space. Our method not only unifies generative and contrastive learning but also extends cross-modal contrastive learning to hyperbolic space. • We propose a comprehensive multi-loss optimization method to tackle the challenges of hierarchical and multi-modal complexities in point cloud representation learning. This method integrates generative, contrastive, and structural learning objectives into a unified design, ensuring a balance between feature diversity and semantic consistency. • We perform thorough qualitative analyses with HyperPoint to demonstrate its potential for capturing hierarchical structures and preserving semantic consistency in cross-modal data. Through extensive experiments, we demonstrate that the hyperbolic counterpart outperforms the Euclidean setting.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

