Return
A Q-learning based path planning method with hyper-heuristic scheduling and escaping based algorithms
DOI:10.1016/j.eswa.2026.132128.png)
Abstract
En 中文
Planning stable and reliable paths for an agent in complex trap environments is a critical challenge in autonomous navigation. Most path planning algorithms suffer from slow convergence, vulnerability to local optima, and the production of suboptimal paths. To address these challenges, this paper proposes an improved Q-learning path planning algorithm (DHH-QL) designed to tackle the difficulties encountered in such environments. The DHH-QL algorithm proposed in this study integrates the evaluation function-based hyper-heuristic scheduling algorithm (QHDF). In QHDF, the evaluation function is used to reasonably schedule various heuristic operators, generating optimal virtual global and escape paths to help the agent quickly escape the trap environment. Additionally, we designed a pruning strategy to address the issue of redundant path generation in the trap environment by pruning through the generation of escape points. Furthermore, DHH-QL introduces an innovative intermittent learning strategy that dynamically adjusts the learning rate based on the agent’s adaptation to the trap environment, enabling adaptive learning. The study was conducted in random environments as well as static and dynamic trap environments, and the results show that DHH-QL outperforms the ACO, GA, traditional QL, O-QL, and TRE-QL algorithms in terms of runtime, path length, smoothness, and convergence speed. The experimental results indicate that DHH-QL, integrating hyper-heuristic scheduling, pruning strategy, and intermittent learning, provides an effective solution for path planning in random and complex trap environments.
Keywords:
Q-learning
path planning
hyper-heuristic scheduling
trap environment
intermittent learning
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

