Return
A hierarchical reinforcement learning-based vehicle-to-grid dispatch architecture for car parks: Integrating proximal policy optimisation and a large language model within a dual-agent framework
J
X
W
F
DOI:10.1016/j.seta.2026.105158.png)
Abstract
En 中文
• A fine-tuned LLM is integrated as a semantic reasoning module to adaptively reweight multi-objective rewards in real time. • A hierarchical dual-agent framework combines PPO-based decision making with convex optimisation-based execution for stable V2G scheduling. • The LLM-driven reweighting mechanism accelerates convergence and enhances adaptability under non-stationary grid conditions. • Natural-language reasoning improves the transparency and interpretability of reward adaptation and policy evolution.
Keywords:
Vehicle-to-grid
Deep reinforcement learning
Large language model
Battery swap station
Hierarchical
Journal
IF:
7
Papers:
4.4K
Citations:
2.2W
