Timeline

Emergent abilities of large language models are described

Jason Wei and co-authors catalogued tasks where accuracy jumped from near-chance to strong performance past a scale threshold, a pattern later disputed as a metric artefact.

  • Ideas & essays
  • Benchmarks & progress
  • Notable

Jason Wei and fifteen co-authors from Google, Stanford, UNC Chapel Hill and DeepMind published “Emergent Abilities of Large Language Models,” a survey defining an emergent ability as one “not present in smaller models but present in larger models” — a capability whose performance could not be predicted by extrapolating the trend from smaller models trained the same way. Across benchmarks including multi-step arithmetic, word unscrambling and several BIG-Bench tasks, the authors reported that model accuracy stayed near random chance until a threshold of scale, then rose sharply, producing the discontinuous jumps in published scaling charts that gave the paper its title.

The paper did not claim to explain why such jumps occurred, and was explicit that the pattern held for some tasks and metrics but not others. Its significance was in naming and cataloguing the phenomenon across many published results at once, at a moment when GPT-3-class models were being scaled further and their capabilities on held-out tasks were the subject of active dispute. The finding fed directly into arguments, on both sides of the AI-risk debate, that further scaling could produce qualitatively new and unpredictable capabilities rather than steady, forecastable improvement.

That reading was challenged the following year. A 2023 paper argued the discontinuities were largely an artefact of the metrics used — that accuracy on many of the same tasks improved smoothly with scale under a continuous metric, and appeared to jump only when scored with a threshold or all-or-nothing measure. The dispute was not fully settled: some tasks showed sharp transitions under any reasonable metric, but the “emergence” framing that had shaped two years of scaling commentary was shown to be, at least in part, a choice of how performance was measured rather than a property of the models themselves.