Person
Stefano Ermon
Appears alongside
Featured in threads
Tracks
- Ideas & essays 2
- Compute & infrastructure 1
Direct Preference Optimization paper reframes RLHF as a classification loss
The method skipped the separate reward model and reinforcement-learning loop, and was later adopted for post-training open models including Zephyr and Tulu.
Ideas & essays
FlashAttention makes exact attention IO-aware
Reordering attention around GPU memory rather than approximating it cut training time and unlocked longer sequences — and became default infrastructure.
Ideas & essays · Compute & infrastructure