arrow
返回

Reward estimation for dialogue policy optimisation

delete2018-09-01
delete8
PRE
AI
P
Pei-Hao Su *
M
Milica Gašić
S
Steve Young
DOI:10.1016/j.csl.2018.02.003delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Viewing dialogue management as a reinforcement learning task enables a system to learn to act optimally by maximising a reward function. This reward function is designed to induce the system behaviour required for the target application and for goal oriented applications, this usually means fulfilling the user's goal as efficiently as possible. However, in real-world spoken dialogue system applications, the reward is hard to measure because the user's goal is frequently known only to the user. Of course, the system can ask the user if the goal has been satisfied but this can be intrusive. Furthermore, in practice, the accuracy of the user's response has been found to be highly variable. This paper presents two approaches to tackling this problem. Firstly, a recurrent neural network is utilised as a task success predictor which is pre-trained from off-line data to estimate task success during subsequent on-line dialogue policy learning. Secondly, an on-line learning framework is described whereby a dialogue policy is jointly trained alongside a reward function modelled as a Gaussian process with active learning. This Gaussian process operates on a fixed dimension embedding which encodes each varying length dialogue. This dialogue embedding is generated in both a supervised and unsupervised fashion using different variants of a recurrent neural network. The experimental results demonstrate the effectiveness of both off-line and on-line methods. These methods enable practical on-line training of dialogue policies in real-world applications. (C) 2018 Published by Elsevier Ltd.
Keyword:
Dialogue systems
Reinforcement learning
Deep learning
Reward estimation
Gaussian process
Active learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

C
Computer Speech and Language
IF:
3.4
论文数:
1.5K
被引数:
2.6K

机构

U
University of Cambridge
学者数:
7.7W
论文数: 7.1W
被引数: 13.7W
引用论文

引用论文

err分享
err收藏
POMDP-based control of workflows for crowdsourcing
err2013-09-01
err79
errOAAI
errDai, Peng; Lin, Christopher H.; Mausam; Weld, Daniel S.
err分享
err收藏
A high temperature variety of BiOF一种高温品种的BiOF
err1983-09-01
err0
PREAI
errSamir Matar; Jean-Maurice Reau; Louis Rabardel; Gérard Demazeau; Paul Hagenmuller
err分享
err收藏
Chondrosarcoma in Norway 1990–2013; an epidemiological and prognostic observational study of a complete national cohort
err2019-01-11
err0
errOAAI
errJoachim Thorkildsen; Ingeborg Taksdal; Bodil Bjerkehagen; Hans Kristian Haugland; Tom Børge Johannesen; Trond Viset; Ole-Jacob Norum; Øyvind Bruland; Olga Zaikova
err分享
err收藏
Radiation damage to DNA–protein specific complexes: estrogen response element–estrogen receptor complex
err2007-01-17
err0
PREAI
errViktorie Štísová; Stephane Goffinont; Melanie Spotheim-Maurizot; Marie Davídková
err分享
err收藏
学者 查看更多内容