arrow
Return

Determining the optimal temperature parameter for Softmax function in reinforcement learning

delete2018-09-01
delete33
PRE
AI
Y
Yulin He
X
Xiaoliang Zhang
W
Wei Ao
J
Joshua Zhexue Huang *
DOI:10.1016/j.asoc.2018.05.012delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The temperature parameter plays an important role in the action selection based on Softmax function which is used to transform an original vector into a probability vector. An efficient method named Opti-Softmax to determine the optimal temperature parameter for Softmax function in reinforcement learning is developed in this paper. Firstly, a new evaluation function is designed to measure the effectiveness of temperature parameter by considering the information-loss of transformation and the diversity among probability vector elements. Secondly, an iterative updating rule is derived to determine the optimal temperature parameter by calculating the minimum of evaluation function. Finally, the experimental results on the synthetic data and D-armed bandit problems demonstrate the feasibility and effectiveness of Opti-Softmax method. (C) 2018 Elsevier B.V. All rights reserved.
Keywords:
Softmax function
Temperature parameter
Probability vector
Reinforcement learning
D-armed bandit problem
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

S
shenzhen university
Scholars:
4.5W
Papers: 3.4W
Citations: 72