arrow
返回

Optimizing multi-domain task-oriented dialogue policy through Sigmoidal Discrete Soft Actor–Critic

delete2026-08-05
delete0
PRE
AI
F
Fatemeh Shamsezat
A
Ali Mohades
S
Saeed Shiry Ghidary *
DOI:10.1016/j.csl.2026.102039delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
• 基于强化学习的面向任务对话策略优化框架。 • 多轮对话系统中的稳定性和鲁棒性提升。 • Soft Actor–Critic (SAC) 在结构化动作空间中的扩展。 • 通过一种新型公式化方法缓解离散 SAC 中的低估偏差。 • 基于 sigmoid 的策略参数化方法以改进探索和决策制定。
Keyword:
Task-oriented dialogue systems
Dialogue policy optimization
Reinforcement learning
Soft actor–critic
Discrete action spaces
Sigmoidal policy
Off-policy learning

期刊

C
Computer Speech and Language
IF:
3.4
论文数:
1.5K
被引数:
2.6K

机构

A
amirkabir university of technology
学者数:
1.1K
论文数: 565
被引数: 0