arrow
返回

Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI

delete2024-11-01
delete0
delete
OA
AI
Y
Yaoxian Song
P
Penglei Sun
H
Haoyu Liu
Z
Zhixu Li *
W
Wei Song *
Y
Yanghua Xiao
X
Xiaofang Zhou
DOI:10.1109/TKDE.2024.3399746delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Embodied AI is one of the most popular studies in artificial intelligence and robotics, which can effectively improve the intelligence of real-world agents (i.e. robots) serving human beings. Scene knowledge is important for an agent to understand the surroundings and make correct decisions in the varied open world. Currently, knowledge base for embodied tasks is missing and most existing work use general knowledge base or pre-trained models to enhance the intelligence of an agent. For conventional knowledge base, it is sparse, insufficient in capacity and cost in data collection. For pre-trained models, they face the uncertainty of knowledge and hard maintenance. To overcome the challenges of scene knowledge, we propose a scene-driven multimodal knowledge graph (Scene-MMKG) construction method combining conventional knowledge engineering and large language models. A unified scene knowledge injection framework is introduced for knowledge representation. To evaluate the advantages of our proposed method, we instantiate Scene-MMKG considering typical indoor robotic functionalities (Manipulation and Mobility), named ManipMob-MMKG. Comparisons in characteristics indicate our instantiated ManipMob-MMKG has broad superiority on data-collection efficiency and knowledge quality. Experimental results on typical embodied tasks show that knowledge-enhanced methods using our instantiated ManipMob-MMKG can improve the performance obviously without re-designing model structures complexly.
Keyword:
Task analysis
Knowledge graphs
Artificial intelligence
Knowledge based systems
Robots
Knowledge engineering
Visualization
Multimodal knowledge graph
scene driven
embodied AI
robotic intelligence

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
Z
Zhejiang Laboratory
学者数:
1.8K
论文数: 1.7K
被引数: 0
引用论文

引用论文

Robots in human environments: Basic autonomous capabilities
err1999-07-01
err147
PREAI
errKhatib, O; Yokoi, K; Brock, O; Chang, K; Casal, A
err分享
err收藏
err分享
err收藏
Never-Ending Learning
err2018-04-24
err394
errOAAI
errMitchell, T.; Cohen, W.; Hruschka, E.; Talukdar, P.; Yang, B.; Betteridge, J.; Carlson, A.; Dalvi, B.; Gardner, M.; Kisiel, B.; Krishnamurthy, J.; Lao, N.; Mazaitis, K.; Mohamed, T.; Nakashole, N.; Platanios, E.; Ritter, A.; Samadi, M.; Settles, B.; Wang, R.; Wijaya, D.; Gupta, A.; Chen, X.; Saparov, A.; Greaves, M.; Welling, J.
err分享
err收藏
DBpedia - A large-scale, multilingual knowledge base extracted from WikipediaDBpedia-从维基百科中提取的大规模多语言知识库
err2015-01-01
err2.0K
errOAAI
errLehmann, Jens; Isele, Robert; Jakob, Max; Jentzsch, Anja; Kontokostas, Dimitris; Mendes, Pablo N.; Hellmann, Sebastian; Morsey, Mohamed; van Kleef, Patrick; Auer, Soeren; Bizer, Christian
err分享
err收藏
Impact of Coastal Hazards on Residents’ Spatial Accessibility to Health Services
err2019-12-01
err0
errOAAI
errGeorgios P. Balomenos; Yujie Hu; Jamie E. Padgett; Kyle Shelton
err分享
err收藏
Targeting the forkhead box protein P1 pathway as a novel therapeutic approach for cardiovascular diseases
err2020-07-09
err0
PREAI
errXin-Ming Liu; Sheng-Li Du; Ran Miao; Le-Feng Wang; Jiu-Chang Zhong
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容