Person
1 entry · 29 May 2023
The method skipped the separate reward model and reinforcement-learning loop, and was later adopted for post-training open models including Zephyr and Tulu.
Ideas & essays