Person
1 entry · 24 May 2022
A single prompt phrase, with no worked examples, lifted GSM8K accuracy from 10.4% to 40.7% — chain-of-thought without the exemplars.
Ideas & essays · Benchmarks & progress