Timeline

Shunyu Yao publishes 'The Second Half', arguing RL environment design now matters more than training

A Princeton researcher and former OpenAI staffer argued the field's bottleneck had shifted from training methods to designing tasks and evaluations that reward real-world usefulness.

  • Ideas & essays
  • Notable

Shunyu Yao, a Princeton researcher who had previously worked on agent and reasoning research at OpenAI, published an essay arguing that AI research had entered a “second half” in which progress would be paced less by new training methods and more by how well the field could define and evaluate the problems it wanted models to solve.

His argument ran roughly as follows: the “first half” of the field, from early machine learning through the language-model era, had been dominated by the search for better training recipes — the winners of that period, he wrote, were “training methods or models, not benchmarks,” because devising a new architecture or optimisation trick was harder and more prestigious than writing a good test set. But once large-scale reinforcement learning “finally works” on top of language pretraining, per Yao, models could apply learned reasoning to almost any task with a well-specified environment, and the remaining differentiator became the quality of the environments and evaluations researchers built. Priors, in his framing, now carried more weight than clever algorithms.

Yao pointed to a gap he called the “utility problem”: models that scored near the top of human ranges on exams and competitive-programming benchmarks had not produced a commensurate change in everyday economic output, in part because standard benchmarks assumed automatic, isolated execution rather than the messier reality of long-running tasks, memory and human collaboration.

The essay circulated widely among researchers working on agents and reinforcement learning and became a common point of reference in 2025 discussions of why labs were investing heavily in RL “environments” — training and evaluation setups built to resemble real tasks — rather than only in larger pretraining runs. It offered a framing, argued from inside the field rather than by outside critics, for a shift already visible in how frontier labs described their roadmaps that year.