arrow
返回

Efficient Q-learning hyperparameter tuning using FOX optimization algorithm

delete2025-03-01
delete0
delete
OA
AI
M
Mahmood A. Jumaah *
Y
Yossra H. Ali
T
Tarik A. Rashid
DOI:10.1016/j.rineng.2025.104341delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Reinforcement learning is a branch of artificial intelligence in which agents learn optimal actions through interactions with their environment. Hyperparameter tuning is crucial for optimizing reinforcement learning algorithms and involves the selection of parameters that can significantly impact learning performance and reward. Conventional Q-learning relies on fixed hyperparameter without tuning throughout the learning process, which is sensitive to the outcomes and can hinder optimal performance. In this paper, a new adaptive hyperparameter tuning method, called Q-learning-FOX (Q-FOX), is proposed. This method utilizes the FOX Optimizer-an optimization algorithm inspired by the hunting behaviour of red foxes-to adaptively optimize the learning rate (alpha) and discount factor (gamma) in the Q-learning. Furthermore, a novel objective function is proposed that maximizes the average Q-values. The FOX utilizes this function to select the optimal solutions with maximum fitness, thereby enhancing the optimization process. The effectiveness of the proposed method is demonstrated through evaluations conducted on two OpenAI Gym control tasks: Cart Pole and Frozen Lake. The proposed method achieved superior cumulative reward compared to established optimization algorithms, as well as fixed and random hyperparameter tuning methods. The fixed and random methods represent the conventional Qlearning. However, the proposed Q-FOX method consistently achieved an average cumulative reward of 500 (the maximum possible) for the Cart Pole task and 0.7389 for the Frozen Lake task across 30 independent runs, demonstrating a 23.37% higher average cumulative reward than conventional Q-learning, which uses established optimization algorithms in both control tasks. Ultimately, the study demonstrates that Q-FOX is superior to tuning hyperparameters adaptively in Q-learning, outperforming established methods.
Keyword:
FOX optimization algorithm
Hyperparameter
Optimization
Q-learning
Reinforcement learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Results in Engineering 封面图
Results in Engineering
IF:
7.9
论文数:
1.1W
被引数:
1.7W

机构

University of Technology - Iraq 封面图
University of Technology - Iraq
学者数:
1.6K
论文数: 1.6K
被引数: 2.7K
U
University of Kurdistan Hewler
学者数:
137
论文数: 140
被引数: 184
引用论文

引用论文

Effects of Acute Aerobic Exercise Versus Acute Zolpidem Intake on Sleep in Individuals with Chronic Insomnia
err2024-06-05
err0
errOAAI
errAriella Rodrigues Cordeiro Rozales; Marcos Gonçalves Santana; Shawn D. Youngstedt; SeungYong Han; Daniela Elias de Assis; Bernardo Pessoa de Assis; Giselle Soares Passos
err分享
err收藏
Grey Wolf Optimizer灰狼优化器
err2014-03-01
err1.3W
PREAI
errMirjalili, Seyedali; Mirjalili, Seyed Mohammad; Lewis, Andrew
err分享
err收藏
Invading with biological weapons: the role of shared disease in ecological invasion
err2008-12-06
err0
PREAI
errSally S. Bell; Andrew White; Jonathan A. Sherratt; Mike Boots
err分享
err收藏
NMR study of copper with vanadium impurities
err1977-07-01
err0
PREAI
errD. M. Follstaedt; C. P. Slichter
err分享
err收藏
Reinforcement learning applications in environmental sustainability: a review强化学习在环境可持续性中的应用: 综述
err2024-03-12
err2
errOAAI
errZuccotto, Maddalena; Castellini, Alberto; La Torre, Davide; Mola, Lapo; Farinelli, Alessandro
err分享
err收藏
学者 查看更多内容