Return
Determining the optimal temperature parameter for Softmax function in reinforcement learning
DOI:10.1016/j.asoc.2018.05.012.png)
Abstract
En 中文
The temperature parameter plays an important role in the action selection based on Softmax function which is used to transform an original vector into a probability vector. An efficient method named Opti-Softmax to determine the optimal temperature parameter for Softmax function in reinforcement learning is developed in this paper. Firstly, a new evaluation function is designed to measure the effectiveness of temperature parameter by considering the information-loss of transformation and the diversity among probability vector elements. Secondly, an iterative updating rule is derived to determine the optimal temperature parameter by calculating the minimum of evaluation function. Finally, the experimental results on the synthetic data and D-armed bandit problems demonstrate the feasibility and effectiveness of Opti-Softmax method. (C) 2018 Elsevier B.V. All rights reserved.
Keywords:
Softmax function
Temperature parameter
Probability vector
Reinforcement learning
D-armed bandit problem
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

