Return
Uncertainty-penalized reinforcement learning from human feedback with diversified reward LoRA ensembles
Y
Y
H
K
D
B
H
DOI:10.1016/j.ipm.2025.104548.png)
Abstract
En 中文
• Fine-tuning large language models to align with human preferences. • An uncertainty-aware reward modeling method utilizing diversified LoRA ensembles. • Uncertainty-penalized reinforcement learning from human feedback to mitigate overoptimization.
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
