Benchmarks · Aggregate indices & arenas
Artificial Analysis Intelligence Index
also: AA Intelligence Index, Artificial Analysis Index
A single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.
Artificial AnalysisReleased Q1 2024Live
Where most benchmarks measure one thing, the Artificial Analysis Intelligence Index tries to compress many into one number. George Cameron and Micah Hill-Smith started the site in 2023 as a side project comparing model price and latency, after Hill-Smith found no independent source tracking which model was actually worth using for a given task. It grew into a continuously updated index that pairs a composite capability score with running price and throughput data pulled directly from providers’ APIs.
The index itself is a moving target by design. Version 4.1.1, current as of August 2026, weights nine separate evaluations into four categories — Agents (34%), Coding (24%), Scientific Reasoning (24%) and General (18%) — drawing on tests such as GDPval-AA, Terminal-Bench and Humanity’s Last Exam. Components and even their grading models are periodically swapped out as older ones saturate, which keeps the index responsive to real capability differences but means scores from different versions cannot be read as one continuous scale. The point was made plainly in September 2026, when version 4.3.2 rescored every model several points lower — GPT-6 Astra fell from 61 to 53 — and Claude Opus 5.5 (max) took the lead at 58, ahead of Claude Fable 5.1 and GPT-6 Astra at 53. Under v4.1.1, Claude Opus 5 (max) had led at 63; earlier in the summer, Zhipu’s open-weight GLM-5.2 had become the highest-scoring open model at 51, a result the company highlighted as evidence openly licensed models were closing on the closed frontier.
Because Artificial Analysis runs its own evaluations rather than relying on developers’ self-reported numbers, its index has become a common reference point in coverage of new releases, cited alongside — and sometimes instead of — a lab’s own benchmark claims. It is not peer-reviewed or academically governed, and its methodology choices, including which evaluations to include or retire, are set unilaterally by the company running it.
The set
A weighted average of Artificial Analysis's own component evaluations. Version 4.1.1 (August 2026) used nine, weighted Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%; version 4.3.2 (September 2026) uses ten, including GDPval-AA v2.1, Terminal-Bench 4.0 and Humanity's Last Exam. Components and grader models are swapped out as older ones saturate, so successive versions are not on a strictly comparable scale.
Where it stands
Actively revised, and versions are not on one scale: v4.3.2 (September 2026) rescored every model several points lower than v4.1.x. Widely cited alongside labs' own benchmark claims, but as a third-party composite rather than an academic publication it is not peer-reviewed.
Editions, and how each was led
AA Intelligence Index v4.3.2current
Current best: Claude Opus 5.5 (max) — 58 Artificial Analysis leaderboard, max effort with fallback. Claude Fable 5.1 53, GPT-6 Astra 53, Claude Opus 5 51, GPT-6 Sol 48, Grok 4.7 46 on the same version.
AA Intelligence Index v4.0–4.1Retired
Current best: Claude Opus 5 (max) — 63 Live leaderboard snapshot under v4.1.1; several models clustered close behind.
In the timeline · 5 entries
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
xAI releases a new Grok model that a benchmark firm rates on par with OpenAI's flagship
Priced the same as its predecessor at $2/$6 per million tokens; Musk said a larger Grok 4.7 was already in training and expected within three to four weeks.
Models & capabilities
South Korea releases a state-backed open model to cut reliance on foreign AI
Motif Technologies built the model from scratch under a South Korean government contest that bars foreign weights, competing to supply a planned national AI assistant.
Open weights & ecosystem · Models & capabilities
Meta launches Muse Spark, its first closed frontier model
Led by former Scale AI chief Alexandr Wang, the model is proprietary and API-only, reversing the open-weight approach Meta had used for the Llama family.
Models & capabilities · Open weights & ecosystem · Labs & people
Artificial Analysis launches independent model benchmarking site
Founded by George Cameron and Micah Hill-Smith as a side project comparing model pricing and latency, it became a widely cited independent reference.
Benchmarks & progress