Threads

The scaling hypothesis and its critics

The bet that more data, parameters and compute keep buying capability — given an empirical spine by GPT-3, revised by Chinchilla, and by 2025 argued to have reached the end of its pretraining era.

The scaling hypothesis is the claim that a neural network’s capability rises predictably with the compute, data and parameters poured into it — and that intelligence might be, in large part, a matter of scale. OpenAI gave it an empirical spine with its 2020 scaling-laws paper, then a demonstration months later with the 175-billion-parameter GPT-3, whose few-shot abilities emerged from size rather than task-specific training.

The hypothesis drew critics as fast as adherents. “On the Dangers of Stochastic Parrots” argued that a larger language model was a better mimic, not a better reasoner, and weighed the costs of building one; Stanford’s foundation-models report named the category even as it questioned it. The recipe was revised from within, too: DeepMind’s Chinchilla paper showed most large models were badly under-trained on data for their size, redrawing the optimal trade-off between parameters and tokens.

GPT-4 was the apex of the pretraining-scale era. Then the axis moved: a test-time-compute paper and OpenAI’s o1 showed that spending more compute at inference — letting a model think for longer — could beat simply building it bigger. By late 2025 even Ilya Sutskever was declaring the scaling era over, arguing the field had to return to new research ideas rather than raw scale. Whether that is the hypothesis failing or merely changing its variable is the argument this thread leaves open.