GAIA, a benchmark for general AI assistants, is released
466 questions that are simple for a person but need browsing, tools and multi-step reasoning to solve; GPT-4 with plugins scored 15% against a 92% human baseline at release.
- Benchmarks & progress
- Notable
Researchers from Meta AI, Hugging Face and the AutoGPT project released GAIA, a benchmark built on an unusual premise: its questions are conceptually simple for a human but hard for an AI system, because answering them requires reasoning, web browsing, tool use and handling several file types rather than recalling memorised facts. The set comprised 466 questions in three difficulty tiers, graded by exact match against a single correct final answer, with most reported scores drawn from a public 165-question validation set.
The gap it exposed was large. On the original questions, human annotators averaged around 92%, while GPT-4 augmented with plugins managed roughly 15% — the paper’s headline evidence that assistants strong on knowledge benchmarks were still weak at the practical, multi-step tasks people actually wanted help with. That framing made GAIA an early standard for the “AI assistant” rather than the chatbot.
As agent scaffolds improved, GAIA scores climbed steeply — open-source Deep Research reproductions and frontier agents pushed validation-set accuracy past 70% within two years — and the leaderboard increasingly measured the surrounding agent system as much as the underlying model. It became one of the reference points against which the computer-use and deep-research agents of 2024–25 were judged.