Return
Sampling as a Structural Constraint for Stable Multitask Offline Reinforcement Learning
DOI:10.3390/app16073511.png)
Abstract
En 中文
Multitask offline reinforcement learning (RL) faces severe instabilities due to heterogeneous data distributions and interference in shared function approximators. Although previous studies address these issues through network architecture modifications, we reinterpret sampling as a structural constraint instead of a performance optimization technique. We propose a two-stage sampling framework. Task-balanced sampling ensures equal task representation in each batch, whereas within-task pairwise ranking maintains relative quality ordering without cross-task value-scale interference. This design promotes stable gradient contributions from the shared function approximators. Through ablation studies on continuous control benchmarks, we demonstrate that removing the pairwise ranking at 20K steps leads to systematic performance degradation across all tasks. Notably, the Hopper task collapses immediately after constraint removal, losing 85% of performance within 1K steps. This demonstrates that pairwise ranking is not a temporary warm-up but a persistent constraint essential throughout training. Our findings establish sampling as a fundamental structural element in multitask offline RL, achieving stability without network architecture modifications.
Keywords:
multitask reinforcement learning
offline reinforcement learning
sampling constraints
pairwise ranking
learning stability
Journal
A
IF:
2.5
Papers:
5.9K
Citations:
4

