Direct Preference Optimization paper reframes RLHF as a classification loss
The method skipped the separate reward model and reinforcement-learning loop, and was later adopted for post-training open models including Zephyr and Tulu.
Stanford HAIIdeas & essays