arrow
Return

Exploring Unknown Environments with Uppaal Stratego: Safe Reinforcement Learning for Navigation and Pump Localization

delete2026-01-01
delete0
PRE
AI
M
Magnus Kallestrup Axelsen
M
Martin Kristjansen *
K
Kim G. Larsen
T
Thomas Grubbe Sandborg Lauritsen
DOI:10.1007/978-3-032-10444-1_15delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As the capabilities and technologies of Unmanned Aerial Vehicles (UAVs) improve, new ways of utilizing them are being investigated. We investigate the use of reinforcement learning to navigate a UAV in an unknown environment, where the room layout is initially unknown. We present a novel approach for exploring and controlling a UAV, which must locate points of interest in such rooms. In this approach, we use reinforcement learning in an online fashion, meaning that the learning is performed multiple times as our knowledge of the room improves. We present the implementation of a stochastic model predictive control approach paired with Q-learning and partition refinement, using Uppaal Stratego to synthesize near-optimal strategies for UAVs to explore, map, and locate objects in environments with no prior knowledge. To ensure the safety of those strategies, we add a pre-shield during learning and employ a post-shield on the proposed actions to be executed. We evaluate our approach using simulation and compare it against a greedy approach, in which the UAV always visits the nearest unexplored part of the map. Our evaluation shows that the approach explores all points of interest approximately 11% faster than the baseline, while also reducing the number of times a new plan must be synthesized by 33%.
Keywords:
Online Reinforcement Learning
Autonomous Drone Navigation
Pre and Post-Shielding
Safe Learning

Journal

S
SOFTWARE ENGINEERING AND FORMAL METHODS, SEFM 2025
IF:
0
Papers:
16
Citations:
0

Organization

A
aalborg university
Scholars:
1.6W
Papers: 1.7W
Citations: 22