arrow
Return

Generalizable Offline Multiobjective Reinforcement Learning via Preference-Conditioned Diffuser

delete2025-10-13
delete0
PRE
AI
Y
Yuchen Xiao
L
Lei Yuan
L
Lihe Li
Z
Ziqian Zhang
Y
Yi-Chen Li
Y
Yang Yu
DOI:10.1109/TNNLS.2025.3591838delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multiobjective reinforcement learning (MORL) addresses sequential decision-making problems with multiple objectives by learning policies optimized for diverse pReferences. While traditional methods necessitate costly online interaction with the environment, recent approaches leverage static datasets containing precollected trajectories, making offline MORL the preferred choice for real-world applications. However, existing offline MORL techniques suffer from limited expressiveness and poor generalization on out-of-distribution (OOD) preferences. To overcome these limitations, we propose diffusion-based MORL (DiffMORL), a generalizable diffusion-based planning frame work for MORL. Leveraging the strong expressiveness and generation capability of diffusion models, DiffMORL further boosts its generalization through offline data mixup, which mitigates the memorization phenomenon and facilitates feature learning by data augmentation. By training on the augmented data, DiffMORL is able to condition on a given preference, whether in-distribution or OOD, to plan the desired trajectory and extract the corresponding action. Evaluations conducted on the datasets for MORL (D4MORL) benchmark demonstrate that DiffMORL achieves state-of-the-art results across nearly all tasks. Notably, it surpasses the best baseline on 14 out of 18 metrics for OOD generalization, underscoring its remarkable generalization ability in offline MORL scenarios.
Keywords:
Diffusion models
generalization
multiobjective reinforcement learning (MORL)
offline reinforcement learning (RL)

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.6K
Citations:
7.2W

Organization

N
nanjing university
Scholars:
7.8W
Papers: 5.6W
Citations: 87
Cited Papers

Cited Papers

Deep Learning for HDD Health Assessment: An Application Based on LSTM
err2022-01-01
err53
PREAI
errDe Santo, Aniello; Galli, Antonio; Gravina, Michela; Moscato, Vincenzo; Sperli, Giancarlo
errShare
errSave
Multi-Objective Neural Evolutionary Algorithm for Combinatorial Optimization Problems
err2023-04-01
err68
PREAI
errShao, Yinan; Lin, Jerry Chun-Wei; Srivastava, Gautam; Guo, Dongdong; Zhang, Hongchun; Yi, Hu; Jolfaei, Alireza
errShare
errSave
U-Net: Convolutional Networks for Biomedical Image Segmentation
err2015-11-18
err0
PREAI
errOlaf Ronneberger; Philipp Fischer; Thomas Brox
errShare
errSave
errShare
errSave
err
IF0
err
err0
errOAAI
err
errShare
errSave
researcher View more