返回
A Replaceable Curiosity-Driven Candidate Agent Exploration Approach for Task-Oriented Dialog Policy Learning
DOI:10.1109/ACCESS.2024.3462719.png)
摘要
En 中文
Task-oriented dialog policy learning is often formulated as a Reinforcement Learning problem whose rewards from the environment are extremely sparse, which means that the agent will often not find the reward by acting randomly. Thus, exploration techniques are of primary importance when solving RL problems, and more sophisticated exploration methods must be devised. In this study, we propose a replaceable curiosity-driven candidate agent exploration approach to encourage the agent to balance action sampling and explore new environments without overly violating dialog strategies. In this framework, we follow the employment of the curiosity model but design weight for the curiosity reward to balance exploration and exploitation. We designed a multi-candidate agent mechanism to filter an agent with relatively balanced action sampling for formal dialog training to motivate agents to escape pseudo-optimal actions in the early training stage. In addition, we propose a replacement mechanism for the first time to prevent the elected agents from performing poorly in the later stages of training and to fully utilize all the candidate agents. The experimental results show that the adjustable curiosity reward promotes dialog policy convergence. The agent replacement mechanism effectively blocks the training of poorly trained agents, significantly increasing the task's average success rate and reducing the number of dialog turns. In this research, an exploration approach for task-oriented dialog system is designed to encourage agents to explore environment through balanced action sampling, without significantly deviating from learned dialog strategies. Compared to baselines, the replaceable curiosity-driven candidate agent exploration approach yields a higher average success rate of 0.714 and a lower number of average turns of 20.6.
Keyword:
Training
Reinforcement learning
Planning
Market research
Optimization
Motion pictures
Indium tin oxide
Multi-agent systems
Dialog management
reinforcement learning
deep Dyna-Q
curiosity
multi-agent optimization
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Random curiosity-driven exploration in deep reinforcement learning深度强化学习中随机好奇心驱动的探索
NEUROCOMPUTING
IF6.5
Cardiovascular disease risk factors and antiretroviral therapy in an HIV‐positive UK population心血管疾病风险因素和抗逆转录病毒疗法在英国HIV阳性人群中的应用
HIV Medicine
IF0
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究

