Benchmarks · Agents, tools & computer use
GAIA
also: GAIA benchmark, General AI Assistants benchmark
Can an AI assistant answer real-world questions that are conceptually simple for a human but require reasoning, web browsing, tool use and handling multiple file types to actually solve?
Meta AI (FAIR), Hugging Face & AutoGPTReleased 21 November 2023Live
GAIA asks a deceptively simple question: can an AI system actually finish the kind of task a competent human assistant would, rather than just talk about it? Released in November 2023 by researchers at Meta, Hugging Face and AutoGPT, its 466 questions are each answerable by a person with internet access and ordinary effort, but require an assistant to chain together web browsing, file reading and multi-step reasoning to reach one exact, checkable answer. The benchmark deliberately inverts the usual difficulty curve: instead of expert-level questions that are hard for humans too, it targets tasks that are easy for humans and hard for models.
That gap was stark at launch — GPT-4 with plugins managed only 15% against an average human score of 92%. The rise of dedicated browsing agents closed much of it: Hugging Face’s open-source reproduction of OpenAI’s Deep Research reached 55% on the validation set within about a year, against roughly 67% the post attributed to OpenAI’s own proprietary system, and by late 2025 Princeton’s independently run HAL leaderboard had general-purpose agents built on Claude Sonnet 4.5 clearing 74%.
GAIA’s scores now say as much about the surrounding agent scaffold — how it browses, retries and calls tools — as about the underlying model: a controlled 2026 study found that changing only the scaffold moved a single model’s GAIA accuracy by as much as 28 points. That is why the numbers split in two. GAIA’s own community-submitted test leaderboard has climbed into the low-90s — ensemble scaffolds such as CustomGPT.ai and Co-Sight Pro report around 93% by 2026, past the 92% human baseline — but the board is open-submission with held answers, and its raw data visibly contains label-probing entries, so its very top scores are plausibly inflated. Standardized independent evaluations tell a more sober story: Princeton’s HAL, which re-runs agents itself on the validation set, plateaus around 75%. GAIA remains one of the standard references for whether an “agent” product actually completes tasks — but which GAIA number is being quoted, and how it was produced, matters as much as the number itself.
The set
466 questions in three difficulty tiers, 300 with held-out answers to keep a live leaderboard; a public 165-question validation set is used for most reported scores. Grading is exact-match against a single correct final answer.
Example
What was the actual enrollment count of the clinical trial on H. pylori in acne vulgaris patients from Jan-May 2018 as listed on the NIH website? (Answer: 90)arxiv.org
Where it stands
Scores depend heavily on the agent scaffold, not the base model alone — a controlled study found scaffold choice can move GAIA accuracy by up to 28 points for a single model. GAIA's own community-submitted TEST leaderboard now shows ensemble scaffolds around 90–93%, but the open-submission format shows signs of overfitting (label-probing entries appear in the raw data), so a standardized independent run — Princeton's HAL, on the validation set — is the cleaner comparison, and there the top sits around 75%.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- November 2023GPT-4 + plugins15%Against a 92% average-human baseline on the same questions, from the original paper.
- February 2025Hugging Face open Deep Research reproduction55.15%Open-source agent built to replicate OpenAI's then-new Deep Research; the same post cites OpenAI's own reported figure of around 67% on the same validation set.
- September 2025HAL Generalist Agent (Claude Sonnet 4.5)74.55%
Current best: HAL Generalist Agent (Claude Sonnet 4.5) — 74.55% Independently run by Princeton's Holistic Agent Leaderboard (HAL) on the validation set — the cleanest model-comparison figure. GAIA's own community TEST leaderboard shows higher numbers (ensemble scaffolds ~90–93% in 2026, e.g. CustomGPT.ai, Co-Sight Pro, OPS-Agentic-Search), but those are multi-model ensembles on an open-submission board with visible overfitting, so they are not comparable model capability. Humans score 92%.