arrow
Return

Fusing differentiable rendering and language-image contrastive learning for superior zero-shot point cloud classification

delete2024-09-01
delete0
PRE
AI
J
Jinlong Xie
L
Long Cheng
G
Gang Wang
M
Min Hu
Z
Zaiyang Yu
M
Minghua Du
宁欣 (Xin Ning) *
DOI:10.1016/j.displa.2024.102773delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-shot point cloud classification involves recognizing categories not encountered during training. Current models often exhibit reduced accuracy on unseen categories without 3D pre-training, emphasizing the need for improved precision and interoperability. We propose a novel approach integrating differentiable rendering with contrastive language-image pre-training. Initially, differentiable rendering autonomously learns representative viewpoints from the data, enabling the transformation of point clouds into multi-view images while preserving key visual information. This transformation facilitates optimized viewpoint selection during training, refining the final feature representation. Features are extracted from the multi-view images and integrated into a global multi-view feature using a cross-attention mechanism. On the textual side, a large language model (LLM) is provided with 3D heuristic prompts to generate 3D-specific text reflecting category-specific traits, from which textual features are derived. The LLM's extensive pre-trained knowledge enables it to capture abstract notions and categorical features relevant to distinct point cloud categories. Visual and textual features are aligned in a unified embedding space, enabling zero-shot classification. Throughout training, the Structural Similarity Index (SSIM) is integrated into the loss function to encourage the model to discern more distinctive viewpoints, reduce redundancy in multi-view imagery, and enhance computational efficiency. Experimental results on the ModelNet10, ModelNet40, and ScanObjectNN datasets demonstrate classification accuracies of 75.68%, 66.42%, and 52.03%, respectively, surpassing prevailing methods in zero-shot point cloud classification accuracy.
Keywords:
Computer vision
Differentiable rendering
Zero-shot learning
Language-image fusion
Point cloud classification

Journal

Displays cover
Displays
IF:
3.4
Papers:
2.2K
Citations:
3.2K

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
N
north china electric power university
Scholars:
2.5W
Papers: 1.7W
Citations: 16
C
chinese people's liberation army general hospital
Scholars:
1.9W
Papers: 1.1W
Citations: 12
I
Imperial College London
Scholars:
8.3W
Papers: 7.3W
Citations: 11.1W
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704
researcher View more organizations