arrow
返回

RPCS v2.0: Object-detection-based recurrent point cloud selection method for 3D dense captioning

delete2024-04-01
delete0
PRE
AI
S
Shinko Hayashi
张智强 封面图
张智强 (Zhang, Zhiqiang)
J
Jinjia Zhou *
DOI:10.1016/j.neucom.2024.127350delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
3D dense captioning is the process of generating natural language descriptions for objects in a 3D scene, represented as RGB-D scans or point clouds. Three problems currently limit the potential performance of existing methods. First, existing methods randomly select points from the point cloud, leading to the exclusion of important points or inclusion of low-value points, for the detected objects. This, in turn, degrades the quality of generated descriptions. Although our previously proposed method, namely Recurrent Point Clouds Selection (RPCS) mitigates aforementioned issue, it possesses an inaccurate termination criterion that causes unexpected interruptions in the loop. Furthermore, existing methods utilize the older object detector, which limits their inherent performance. Finally, the descriptions generated by existing methods only describe individual detected objects in the scene, which may be inconvenient for viewers from a visual perspective. To address these problems, we propose RPCS v2.0, which features several improvements over our original design. To maintain a high quality of generated descriptions while avoiding the unexpected interruptions inherent to the original RPCS method, we propose a modified termination criterion that continuously compares the element counts of the bad list. To address limitations associated with the older object detector, we implemented a newer object detector to further improve performance. Finally, to ensure the user-friendliness of visuals, we leveraged a language model to summarize the generated descriptions, producing a comprehensive representation of the entire scene. Our proposed approach, referred to as ScanT3D, significantly mitigated the issue of suboptimal description generation and outperformed the state-of-the-art methods by a large margin (65.84%CiDEr@0.5IoU).
Keyword:
3D dense captioning
Point clouds
Recurrent point clouds selection

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

H
Hosei University
学者数:
953
论文数: 1.1K
被引数: 718
引用论文

引用论文

Fast and optimal generation of entanglement in bosonic Josephson junctions
err2019-02-25
err0
errOAAI
errGiacomo Sorelli; Manuel Gessner; Augusto Smerzi; Luca Pezzè
err分享
err收藏
Interactions Guided Generative Adversarial Network for unsupervised image captioning
err2020-12-01
err23
PREAI
errCao, Shan; An, Gaoyun; Zheng, Zhenxing; Ruan, Qiuqi
err分享
err收藏
Relation constraint self-attention for image captioning图像字幕的关系约束自注意
err2022-08-01
err16
PREAI
errJi, Junzhong; Wang, Mingzhan; Zhang, Xiaodan; Lei, Minglong; Qu, Liangqiong
err分享
err收藏
Estimation of parameters of various damping models in planar motion of a pendulum
err2020-07-07
err0
errOAAI
errRobert Salamon; Henryk Kamiński; Paweł Fritzkowski
err分享
err收藏
Multidrug Resistance Gene Expression in Childhood Medulloblastoma: Correlation with Clinical Outcome and DNA Ploidy in 29 Patients
err1995-01-01
err0
PREAI
errPauline M. Chou; Miguel Reyes-Mugica; Nora Barquin; Takasumi Yasuda; Xiaodi Tan; Tadanori Tomita
err分享
err收藏
学者 查看更多内容