arrow
返回

Indoor Scene Change Captioning Based on Multimodality Data

delete2020-08-23
delete16
delete
OA
AI
Y
Yue Qiu *
Y
Yutaka Satoh
R
Ryota Suzuki
K
Kenji Iwata
K
Kataoka, Hirokatsu
DOI:10.3390/s20174761delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This study proposes a framework for describing a scene change using natural language text based on indoor scene observations conducted before and after a scene change. The recognition of scene changes plays an essential role in a variety of real-world applications, such as scene anomaly detection. Most scene understanding research has focused on static scenes. Most existing scene change captioning methods detect scene changes from single-view RGB images, neglecting the underlying three-dimensional structures. Previous three-dimensional scene change captioning methods use simulated scenes consisting of geometry primitives, making it unsuitable for real-world applications. To solve these problems, we automatically generated large-scale indoor scene change caption datasets. We propose an end-to-end framework for describing scene changes from various input modalities, namely, RGB images, depth images, and point cloud data, which are available in most robot applications. We conducted experiments with various input modalities and models and evaluated model performance using datasets with various levels of complexity. Experimental results show that the models that combine RGB images and point cloud data as input achieve high performance in sentence generation and caption correctness and are robust for change type understanding for datasets with high complexity. The developed datasets and models contribute to the study of indoor scene change understanding.
Keyword:
image captioning
three-dimensional (3D) vision
deep learning
human-robot interaction
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Sensors 封面图
Sensors
IF:
3.5
论文数:
7.2W
被引数:
20.9W

机构

U
University of Tsukuba
学者数:
1.8W
论文数: 1.5W
被引数: 1.7W
引用论文

引用论文

End-to-End Multimodal Emotion Recognition Using Deep Neural Networks基于深度神经网络的端到端多模态情感识别
err2017-12-01
err375
errOAAI
errTzirakis, Panagiotis; Trigeorgis, George; Nicolaou, Mihalis A.; Schuller, Bjorn W.; Zafeiriou, Stefanos
err分享
err收藏
Synthesis of advanced fluorescent probes — water-soluble symmetrical tricarbocyanines with phosphonate groups
err2017-05-17
err0
PREAI
errT. A. Podrugina; V. V. Temnov; I. A. Doroshenko; V. A. Kuzmin; T. D. Nekipelova; M. V. Proskurnina; N. S. Zefirov
err分享
err收藏
On the Existence of LiFeVO4 – Tales and Imagination
err2011-04-12
err0
PREAI
errOliver Clemens; Matthias Bauer; Robert Haberkorn; Horst Philipp Beck
err分享
err收藏
学者 查看更多内容