RLHF (Reinforcement Learning from Human Feedback)
A training technique that aligns models with human preferences using a reward model learned from human ratings.
A training technique that aligns models with human preferences using a reward model learned from human ratings.