FlashAttention makes exact attention IO-aware
Reordering attention around GPU memory rather than approximating it cut training time and unlocked longer sequences — and became default infrastructure.
Stanford HAIIdeas & essays · Compute & infrastructure