Return
VATE: Variational Attention Trajectory Encoder for Preference-Based Reinforcement Learning
DOI:10.1109/LSP.2026.3697690.png)
Abstract
En 中文
Preference-based reinforcement learning (PbRL) enables agents to learn from human feedback without explicit reward engineering. However, existing methods rely on simple MLP architectures that treat state-action pairs independently, failing to exploit the temporal structure inherent in behavioral signals. Preference learning can be cast as estimating an underlying utility signal from sparse, noisy binary observations. We propose VATE (Variational Attention Trajectory Encoder), an online PbRL framework that integrates: (1) a variational encoder for robust signal manifold learning, (2) transformer-based temporal modeling for long-range dependency capture, and (3) multi-scale attention aggregation for adaptive signal fusion. Experiments on Meta-World and DMControl tasks demonstrate that VATE achieves strong sample efficiency and robustness over state-of-the-art baselines; additional teacher-sensitivity and ablation studies support its robustness under noisy feedback and validate each core component.
Keywords:
Preference-based reinforcement learning
utility signal estimation
variational representation learning
multi-scale attention aggregation

