返回
Task-Aware Semantic Map++: Cost-Efficient Task Assignment With Advanced Benchmark
DOI:10.1109/LRA.2026.3656794.png)
摘要
En 中文
Enabling robots to perform diverse tasks autonomously requires a sophisticated semantic understanding of 3D scenes. However, conventional scene representations, which primarily rely on static attributes like visual information or object labels, have significant limitations in allowing robots to infer context-aware actions. We introduce Task-Aware Semantic Map++ (TASMap++), the framework that overcomes these limitations by constructing a map that assigns appropriate tasks to objects based on their holistic context. While prior work like TASMap pioneered this task-centric approach, it suffered from high computational costs and inaccuracies due to its reliance on single-frame analysis, which often fails to capture an object’s complete state. In contrast, TASMap++ resolves these issues with a multi-view synthesis pipeline that integrates multiple perspectives of an object for task assignment, resulting in significantly improved computational efficiency over its predecessor. Furthermore, to overcome biases in the existing TASMap evaluation, we established a reliable benchmark derived from the consensus of 32 participants across 231 cluttered scenes. On this benchmark, TASMap++ demonstrates superior accuracy over baselines. Finally, we introduce context-aware grounding, a paradigm distinct from conventional object grounding that relies on visual and spatial attributes. We present a downstream application of TASMap++ as a method to address this challenge and show experimentally that conventional grounding methods struggle in this setting, whereas TASMap++ is markedly more effective. To confirm these findings, the framework’s robustness and practicality were validated through extensive experiments on 3D indoor datasets, including real-world scan datasets.
Keyword:
Semantic scene understanding
mapping
AI-based methods
期刊
I
IF:
5.3
论文数:
1.7K
被引数:
3.9W
机构
引用论文
Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal NavigationHabitat Synthetic Scenes Dataset (HSSD-200):关于对象目标导航中3D场景尺度与真实感权衡的分析
Y. Zhang, P. Yan, H. Tang, J. Zhang, Rapid detection of tear lactoferrin for diagnosis of dry eyes by using fluorescence polarization-based aptasensor, Scientific Reports, 13 (2023) 15179. DOI: 10.1038/s41598-023-42484-5.Y. Zhang, P. Yan, H. Tang, J. Zhang, 基于荧光偏振的适体传感器用于快速检测泪液乳铁蛋白以诊断干眼症, Scientific Reports, 13 (2023) 15179. DOI: 10.1038/s41598-023-42484-5.

