Return
Personalized Dialogue Policy Learning Framework Based on Implicit User Profiles
DOI:10.1109/THMS.2025.3620160.png)
Abstract
En 中文
Dialogue policy is a core module in pipeline dialogue systems as it drives conversation generation. Personalized dialogue policies aim to equip chatbots with tailored personalities, making them behave like real users, providing more accurate action responses, and improving the anthropomorphic capabilities of personal assistants. Yet existing dialogue policy approaches often overlook individual personalities because obtaining explicit user profiles is costly and time-consuming. In this article, we propose a personalized dialogue policy learning framework, named PDL. It dynamically learns implicit user profiles from successful dialogue trajectories. Specifically, we collect a lot of success histories from human–computer interactions to extract sequences of user belief states and agent actions. The extracted sequences are processed via a loop clipping operation and modeled with an autoregressive transformer to mimic human analytical behavior. After that, a new user’s latent personalized preferences are predicted based on the autoregressive transformer model. The personalized preferences are employed to implement dialogue policies via three categories of reinforcement learning algorithms, including value-based approaches, policy-based approaches, and model-based approaches. The experiments are conducted on three different task-oriented dialogue datasets, and the results show that the proposed PDL framework achieves state-of-the-art results compared to other comparative approaches.
Keywords:
Deep reinforcement learning
dialogue policy learning
human–agent interaction
personalized agent
Journal
IF:
4.4
Papers:
1.1K
Citations:
3.5K

