Benchmarks · Aggregate indices & arenas

Artificial Analysis Intelligence Index

also: AA Intelligence Index, Artificial Analysis Index

A single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.

Artificial AnalysisReleased Q1 2024Live

Where most benchmarks measure one thing, the Artificial Analysis Intelligence Index tries to compress many into one number. George Cameron and Micah Hill-Smith started the site in 2023 as a side project comparing model price and latency, after Hill-Smith found no independent source tracking which model was actually worth using for a given task. It grew into a continuously updated index that pairs a composite capability score with running price and throughput data pulled directly from providers’ APIs.

The index itself is a moving target by design. Version 4.1.1, current as of August 2026, weights nine separate evaluations into four categories — Agents (34%), Coding (24%), Scientific Reasoning (24%) and General (18%) — drawing on tests such as GDPval-AA, Terminal-Bench and Humanity’s Last Exam. Components and even their grading models are periodically swapped out as older ones saturate, which keeps the index responsive to real capability differences but means scores from different versions cannot be read as one continuous scale. The point was made plainly in September 2026, when version 4.3.2 rescored every model several points lower — GPT-6 Astra fell from 61 to 53 — and Claude Opus 5.5 (max) took the lead at 58, ahead of Claude Fable 5.1 and GPT-6 Astra at 53. Under v4.1.1, Claude Opus 5 (max) had led at 63; earlier in the summer, Zhipu’s open-weight GLM-5.2 had become the highest-scoring open model at 51, a result the company highlighted as evidence openly licensed models were closing on the closed frontier.

Because Artificial Analysis runs its own evaluations rather than relying on developers’ self-reported numbers, its index has become a common reference point in coverage of new releases, cited alongside — and sometimes instead of — a lab’s own benchmark claims. It is not peer-reviewed or academically governed, and its methodology choices, including which evaluations to include or retire, are set unilaterally by the company running it.

The set

A weighted average of Artificial Analysis's own component evaluations. Version 4.1.1 (August 2026) used nine, weighted Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%; version 4.3.2 (September 2026) uses ten, including GDPval-AA v2.1, Terminal-Bench 4.0 and Humanity's Last Exam. Components and grader models are swapped out as older ones saturate, so successive versions are not on a strictly comparable scale.

Where it stands

Actively revised, and versions are not on one scale: v4.3.2 (September 2026) rescored every model several points lower than v4.1.x. Widely cited alongside labs' own benchmark claims, but as a third-party composite rather than an academic publication it is not peer-reviewed.

Editions, and how each was led

AA Intelligence Index v4.3.2current

Released September 2026Ten components: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Scores run several points below v4.1.x for the same model (GPT-6 Astra 61 → 53).

  1. September 2026Claude Opus 5.5 (max)58 (v4.3.2)Top of v4.3.2 at release, five points clear of Fable 5.1 and GPT-6 Astra.

Current best: Claude Opus 5.5 (max) — 58 Artificial Analysis leaderboard, max effort with fallback. Claude Fable 5.1 53, GPT-6 Astra 53, Claude Opus 5 51, GPT-6 Sol 48, Grok 4.7 46 on the same version.

AA Intelligence Index v4.0–4.1Retired

Released June 2026Nine components incl. GDPval-AA, τ³-Banking, Terminal-Bench v2.1, SciCode, HLE, GPQA Diamond, CritPt, AA-LCR and AA-Omniscience. The frontier's scores here were gathered June–September 2026 across point releases up to v4.1.1, so they are only loosely comparable with one another.

  1. June 2026GLM-5.251 (highest open-weight score at the time)Zhipu's MIT-licensed model topped the open-weight segment of the index rather than the overall leaderboard.

Current best: Claude Opus 5 (max) — 63 Live leaderboard snapshot under v4.1.1; several models clustered close behind.

In the timeline · 5 entries

More aggregate indices & arenas benchmarks