arrow
返回

A Replaceable Curiosity-Driven Candidate Agent Exploration Approach for Task-Oriented Dialog Policy Learning

delete2024-01-01
delete0
delete
OA
AI
X
Xuecheng Niu *
A
Akinori Ito
T
Takashi Nose
DOI:10.1109/ACCESS.2024.3462719delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Task-oriented dialog policy learning is often formulated as a Reinforcement Learning problem whose rewards from the environment are extremely sparse, which means that the agent will often not find the reward by acting randomly. Thus, exploration techniques are of primary importance when solving RL problems, and more sophisticated exploration methods must be devised. In this study, we propose a replaceable curiosity-driven candidate agent exploration approach to encourage the agent to balance action sampling and explore new environments without overly violating dialog strategies. In this framework, we follow the employment of the curiosity model but design weight for the curiosity reward to balance exploration and exploitation. We designed a multi-candidate agent mechanism to filter an agent with relatively balanced action sampling for formal dialog training to motivate agents to escape pseudo-optimal actions in the early training stage. In addition, we propose a replacement mechanism for the first time to prevent the elected agents from performing poorly in the later stages of training and to fully utilize all the candidate agents. The experimental results show that the adjustable curiosity reward promotes dialog policy convergence. The agent replacement mechanism effectively blocks the training of poorly trained agents, significantly increasing the task's average success rate and reducing the number of dialog turns. In this research, an exploration approach for task-oriented dialog system is designed to encourage agents to explore environment through balanced action sampling, without significantly deviating from learned dialog strategies. Compared to baselines, the replaceable curiosity-driven candidate agent exploration approach yields a higher average success rate of 0.714 and a lower number of average turns of 20.6.
Keyword:
Training
Reinforcement learning
Planning
Market research
Optimization
Motion pictures
Indium tin oxide
Multi-agent systems
Dialog management
reinforcement learning
deep Dyna-Q
curiosity
multi-agent optimization

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

T
tohoku university
学者数:
4.3W
论文数: 3.6W
被引数: 31
引用论文

引用论文

err分享
err收藏
R2* Impact on Hepatic Fat Quantification With a Commercial Single Voxel Technique at 1.5 and 3.0 T
err2024-06-04
err0
errOAAI
errVéronique Fortier; Ahmed Mohamed; Evan McNabb; Jérémy Dana; Rita Zakarian; Ives R. Levesque; Caroline Reinhold
err分享
err收藏
Reinforcement Learning for Mobile Robotics Exploration: A Survey
err2023-08-01
err50
PREAI
errGaraffa, Luiza Caetano; Basso, Maik; Konzen, Andrea Aparecida; de Freitas, Edison Pignaton
err分享
err收藏
err分享
err收藏
Exploration in deep reinforcement learning: A survey深度强化学习探索: 综述
err2022-09-01
err138
errOAAI
errLadosz, Pawel; Weng, Lilian; Kim, Minwoo; Oh, Hyondong
err分享
err收藏
Corticosteroids Do Not Influence the Efficacy and Kinetics of CAR-T Cells for B-Cell Acute Lymphoblastic Leukemia
err2019-11-13
err0
errOAAI
errShuangyou Liu; Biping Deng; Jing PAN; Zhichao Yin; Yuehui Lin; Zhuojun Ling; Tong Wu; Zhiyong Gao; Yanzhi Song; Yongqiang Zhao; Chunrong Tong
err分享
err收藏
学者 查看更多内容