arrow
返回

An Improved Multi-View Convolutional Neural Network for 3D Object Retrieval

delete2020-01-01
delete16
PRE
AI
X
Xinwei He
S
Song Bai *
J
Jiajia Chu
白
白翔 (Xiang Bai)
DOI:10.1109/TIP.2020.3008970delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Learning robust and discriminative representations is essential for 3D object retrieval. In this paper, we present an improved Multi-view Convolutional Neural Network (MVCNN) for view-based 3D object representation learning. Our technical contributions are divided into two aspects. First, we propose to employ Group-view Similarity Learning (GSL) over the multi-view representations before the aggregation operation (i.e., max-pooling in MVCNN). We assume that the similarity information among the view groups of different 3D objects can provide an important cue but has been neglected more or less by previous methods. To enhance it, we add a branch to the original MVCNN architecture and learn to maintain such group-view similarity relationships. Second, we utilize an end-to-end metric learning loss function to improve the representation learning process. In particular, we propose an improved Triplet-Center Loss (TCL) named Adaptive Margin based Triplet-Center Loss (AMTCL). The original TCL assumes a fixed and common margin to control the relative distance relationship between a sample to its corresponding class center and to the nearest negative center. Though TCL has demonstrated its great capacity on the 3D object retrieval task, however, when considering the distinguishability between samples of one class and samples of another class, we assume that it would be more appropriate that the margin takes different values based on the distinguishability of samples of different classes. Therefore we propose to adaptively and dynamically adjust the margin hyperparameter based on the normalized confusion matrix which is obtained on the training set during the training process. Extensive experiments on several public 3D shape benchmarks show that our method, GSL + AMTCL, can learn more suitable representations for 3D object retrieval, obtaining superior performance against state-of-the-art methods.
Keyword:
Three-dimensional displays
Shape
Feature extraction
Measurement
Training
Convolutional neural networks
Task analysis
3D object retrieval
metric learning
similarity learning
loss functions
multi-view CNN
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

暂无机构信息
引用论文

引用论文

err分享
err收藏
Learning View-Model Joint Relevance for 3D Object Retrieval
err2015-05-01
err43
errOAAI
errLu, Ke; He, Ning; Xue, Jian; Dong, Jiyang; Shao, Ling
err分享
err收藏
err分享
err收藏
err分享
err收藏
Automatic Instrument Segmentation in Robot-Assisted Surgery Using Deep Learning
err
IF0
err2018-03-03
err0
errOAAI
errAlexey A. Shvets; Alexander Rakhlin; Alexandr A. Kalinin; Vladimir I. Iglovikov
err分享
err收藏
学者 查看更多内容