arrow
Return

Diffusion Dynamic Model for Unsupervised Reinforcement Learning

delete2026-01-01
delete0
PRE
AI
R
Ran Chen
胡小亮 cover
胡小亮 (Xiaoliang Hu)
崔振 (Zhen Cui)
L
Luying Wu
T
Tong Zhang *
DOI:10.1007/978-981-95-5693-9_1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Unsupervised reinforcement learning (URL) enables agents to adapt efficiently to novel tasks but relies on robust environmental modeling. Traditional world models enhance exploration but are prone to single-step errors, leading to distributional shifts. To address this challenge, we introduce the Diffusion Dynamic Model (DDM), incorporating diffusion models into the environmental modeling component of unsupervised reinforcement learning. During pre-training, DDM generates future observations conditioned on current observations and actions. These generated observations are used to design an intrinsic reward function that incentivizes the agent to explore diverse and high-uncertainty states. In the fine-tuning phase, DDM leverages its generative capabilities to perform data augmentation. By producing diverse, high-quality synthetic data, DDM expands the training dataset, enabling the agent to generalize better and adapt more rapidly to task-specific environments. Our approach is validated across three domains and twelve downstream tasks in the URLB benchmark, demonstrating superior exploration and adaptability compared to existing methods.
Keywords:
Diffusion dynamic model
Unsupervised reinforcement learning
Data augmentation

Journal

P
PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2025, PT II
IF:
0
Papers:
27
Citations:
0

Organization

N
nanjing university of science & technology
Scholars:
1.6K
Papers: 534
Citations: 0
B
beijing normal university
Scholars:
4.8K
Papers: 2.0K
Citations: 0