StartupBench finds top agents finish only a third of real startup tasks
Tasks came from paying-customer workflows at AI startups rather than researcher-chosen problems; the strongest general-purpose agent completed about 30% of them end-to-end.
- Benchmarks & progress
- Models & capabilities
- Minor
A 38-author team, led by contributors including Jingzhe Ding and drawing on affiliations that the arXiv listing does not itemise by name, published StartupBench, an agent benchmark built from workflows drawn from AI startup products with real paying users, rather than tasks a researcher selected for the purpose. The authors argued that existing agent benchmarks “largely rely on researcher-selected tasks,” leaving open how well reported scores predict performance on the kind of work a startup actually sells.
Testing general-purpose agents against these market-validated workflows, the authors reported that even the strongest model completed only about 30% of StartupBench end-to-end. They attributed the shortfall chiefly to complex, multi-step instruction following and to gaps in domain-specific expertise, rather than to any single capability current agents uniformly lack.
The benchmark adds to a run of 2026 evaluations built specifically to counter saturation on established agent tests, following efforts such as UC Berkeley’s “Agents’ Last Exam,” which similarly found frontier agents clearing well under half of professional-grade tasks. Its distinguishing claim is provenance: because the tasks are drawn from products with existing commercial demand, a low completion rate is harder to dismiss as an artefact of an unrealistic academic test.