Return
Efficient Reinforcement Learning from Human Feedback via Bayesian preference inference
DOI:10.1016/j.ifacsc.2026.100398.png)
Abstract
En 中文
Learning from human preferences is essential for aligning machine learning models with subjective judgments, but collecting preference data is costly. We study a hybrid framework that combines the scalability of Reinforcement Learning from Human Feedback (RLHF), which trains neural reward models from pairwise comparisons, with the sample efficiency of Preferential Bayesian Optimization. The method integrates Laplace-based Bayesian uncertainty estimation to guide informative preference queries. On high-dimensional Rosenbrock optimization, the approach successfully converges in problems with up to 50 dimensions. In large language model (LLM) fine-tuning, it improves reward-model accuracy by a value within 6-14% under limited annotation budgets. (c) 2026 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Keywords:
Human-in-the-Loop optimization
Reinforcement Learning from Human
Feedback (RLHF)
Preferential Bayesian Optimization (PBO)
Active learning
Preference-based optimization
Large language models (LLMs)
High-dimensional optimization
Journal
I
IF:
1.8
Papers:
80
Citations:
317

