Return
Think fast and far: Long-horizon online POMDP planning via rapid state sampling
Y
E
J
W
Z
L
H
DOI:10.1177/02783649261453546.png)
Abstract
En 中文
<jats:p>
Partially observable Markov decision processes (
<jats:sc>pomdp</jats:sc>
s) are a general and principled framework for motion planning under uncertainty. Despite tremendous improvement in the scalability of
<jats:sc>pomdp</jats:sc>
solvers, long-horizon
<jats:sc>pomdp</jats:sc>
s remain difficult to solve. To alleviate the difficulty, this paper proposes a new approximate online
<jats:sc>pomdp</jats:sc>
solver, called reference-based online
<jats:sc>pomdp</jats:sc>
planning via rapid state space sampling (
<jats:sc>rop-ras3</jats:sc>
).
<jats:sc>rop-ras3</jats:sc>
uses novel extremely fast sampling-based motion planning techniques to sample the state space and generate a diverse set of macro-actions online, which are then used to bias belief-space sampling and infer high-quality policies
<jats:italic toggle="yes">without</jats:italic>
requiring exhaustive enumeration of the action space—a fundamental constraint for modern online
<jats:sc>pomdp</jats:sc>
solvers.
<jats:sc>rop-ras3</jats:sc>
converges to a near-optimal reference-based solution at a rate that depends on the number of sampled actions, rather than the size of the action space.
<jats:sc>rop-ras3</jats:sc>
is evaluated on various long-horizon
<jats:sc>pomdp</jats:sc>
s with up to 3000 lookahead steps and 35-dimensional state spaces, where the state, action and observation spaces can be continuous, discrete, or a hybrid of discrete and continuous. Although the reference-based optimal solution may not be the same as the optimal
<jats:sc>pomdp</jats:sc>
solution, empirical results indicate that in all of these problems, in terms of success rate,
<jats:sc>rop-ras3</jats:sc>
<jats:italic toggle="yes">outperforms</jats:italic>
other state-of-the-art methods by up to
<jats:italic toggle="yes">multiple folds</jats:italic>
. We also demonstrate the capability of our approach on a physical robot demonstration. This work extends the theory and empirical results of our ISRR24 paper. Code can be found at
<jats:ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/RDLLab/ROPRAS3">https://github.com/RDLLab/ROPRAS3</jats:ext-link>
.
</jats:p>
Journal
IF:
5
Papers:
2.4K
Citations:
1.5W
