返回
A Multi-Agent Approach to Modeling Task-Oriented Dialog Policy Learning
DOI:10.1109/ACCESS.2025.3529469.png)
摘要
En 中文
Dialogue policy is a critical research area in human-computer interaction, vital for guiding dialogue generation and improving controllability and interpretability. Multi-agent dialogue policy learning demonstrates superior learning speed and exploration capabilities, positioning it as a promising approach for developing more effective and adaptive dialogue agents. However, many studies neglect to holistically model collaboration between agents, which limits the effectiveness of policy learning. Therefore, this paper proposes a new multi-agent group collaboration mechanism for dialogue policy learning, named GMPL. Concretely, we employ an Actor-Critic network to implement the proposed model, alternately updating individual dialogue agents to optimize policy selection. In each update, we utilize the maximum action value function to determine the appropriate dialogue action, while the maximum state value function serves to guide the policy learning process. This integrated approach ensures that both decision-making and learning phases are effectively aligned, thereby enhancing the overall performance of the dialogue agents. Furthermore, we conduct a theoretical analysis of the convergence properties of the proposed model. Experiments were conducted on two distinct task-oriented dialogue datasets, revealing that the proposed multi-agent model exhibits a significantly faster learning speed and a higher dialogue success rate compared to baseline approaches.
Keyword:
Natural language generation
Reinforcement learning
Decision making
Collaboration
Pipelines
Medical diagnosis
Human computer interaction
Databases
Convergence
Multi-agent systems
Human-computer interaction
dialogue policy learning
deep reinforcement learning
multi-agent learning
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Acquiring New Knowledge Without Losing Old Ones for Effective Continual Dialogue Policy Learning在不丢失旧知识的情况下获取新知识,以进行有效的持续对话政策学习
Cucker-Smale Flocking Behavior for Multiagent Networks With Coopetition Interactions and Communication Delays具有合作竞争交互和通信延迟的多代理网络的cucker-smale蜂拥行为
Multi-Agent Reinforcement Learning for Intelligent V2G Integration in Future Transportation Systems未来交通系统中智能V2G集成的多智能体强化学习
A Survey on Recent Advances and Challenges in Reinforcement Learning Methods for Task-oriented Dialogue Policy Learning任务型对话政策学习强化学习方法的最新进展与挑战综述

