返回
UAV Coverage Path Planning With Quantum-Based Recurrent Deep Deterministic Policy Gradient
DOI:10.1109/TVT.2023.3347219.png)
摘要
En 中文
This study proposes quantum-based deep deterministic policy gradient (Q-DDPG) and quantum-based recurrent DDPG (Q-RDDPG) schemes for time-series optimization in UAV communications. Herein, Q-DDPG-based actor-critic reinforcement learning is utilized to optimize action selections in a large state and continuous action space. In this scheme, quantum models are exploited to reduce computational complexity and training loss. As a particular case, Q-DDPG and Q-RDDPG are employed for trajectory optimization and dynamic resource allocation in UAV communications. The results demonstrate that Q-DDPG and Q-RDDPG schemes achieved higher rewards with lower training losses compared to classical DDPG.
Keyword:
Autonomous aerial vehicles
Training
Optimization
NOMA
Encoding
Vehicle dynamics
Resource management
Deep deterministic policy gradient
energy efficiency
quantum embedding
recurrent
UAV communications
期刊
IF:
7.1
论文数:
1.8W
被引数:
6.6W
机构
引用论文
Spectrum-Sharing UAV-Assisted Mission-Critical Communication: Learning-Aided Real-Time Optimisation
IEEE ACCESS
IF3.6
Deep Deterministic Policy Gradient (DDPG)-Based Resource Allocation Scheme for NOMA Vehicular Communications
IEEE ACCESS
IF3.6
Futuristic view of the Internet of Quantum Drones: Review, challenges and research agenda量子无人机互联网的未来观点: 回顾、挑战和研究议程
When Entanglement Meets Classical Communications: Quantum Teleportation for the Quantum Internet当纠缠遇到经典通信时: 量子互联网的量子隐形传态

