An interactive radar · 58 models · 40 benchmarks · 579 sourced scores

The Race for the Frontier

Every major AI model against the benchmarks that defined the field, drawn as a radar (the same data is also on the line-chart view). Every benchmark is a spoke from the centre (easiest near the top, hardest opposite); a model reaches further out the better it scores, so each model is a polygon and the best-on-every-benchmark frontier is the outer hull. Every axis is normalised to its own 0–100%. Drag the time slider to watch the shape grow; hover a model to isolate it.

now

Normalised per benchmark (0–100% of each axis's own range). A model's polygon only touches the benchmarks it has a sourced score on. Hollow points are company-reported rather than independently confirmed. The full sourced-data matrix is in the table below; the same data is also on the line-chart view.

Every score, in full

The complete matrix behind the chart — 579 sourced scores across 40 benchmarks (rows, oldest and most-saturated first) and 58 models (columns, by release date), so you can see exactly what we have and where the gaps are: 25% of possible cells are filled. Italic scores are company-reported rather than independently confirmed. Hover a score for the exact setting and provenance; click it to open the source. Scroll the table both ways.

Benchmark ↓   Model →GPT-4OpenAI · March 2023Claude 2Anthropic · July 2023Claude 3 OpusAnthropic · March 2024GPT-4oOpenAI · May 2024Claude 3.5 SonnetAnthropic · June 2024Llama 3.1 405BMeta AI · July 2024o1-previewOpenAI · September 2024o1OpenAI · December 2024o3OpenAI · December 2024DeepSeek-R1DeepSeek · January 2025Gemini 2.5 ProGoogle DeepMind · March 2025Qwen3 235BAlibaba / Qwen · April 2025Claude Opus 4Anthropic · May 2025Grok 4xAI · July 2025Kimi K2Moonshot AI · July 2025GPT-5OpenAI · August 2025Qwen3-MaxAlibaba / Qwen · September 2025Claude Sonnet 4.5Anthropic · September 2025GPT-5.1OpenAI · October 2025Gemini 3 ProGoogle DeepMind · November 2025Claude Opus 4.5Anthropic · November 2025GPT-5.2OpenAI · December 2025Gemini 3 FlashGoogle DeepMind · December 2025Claude Opus 4.6Anthropic · February 2026GLM-5Zhipu AI · February 2026Gemini 3 Deep Think v2Google DeepMind · February 2026Claude Sonnet 4.6Anthropic · February 2026Gemini 3.1 ProGoogle DeepMind · February 2026GPT-5.4OpenAI · March 2026Claude Mythos PreviewAnthropic · April 2026Claude Opus 4.7Anthropic · April 2026Kimi K2.6Moonshot AI · April 2026GPT-5.5OpenAI · April 2026DeepSeek V4DeepSeek · April 2026Gemini 3.5 FlashGoogle DeepMind · May 2026Claude Opus 4.8Anthropic · May 2026Muse Spark 1.1Muse · June 2026Claude Fable 5Anthropic · June 2026Claude Mythos 5Anthropic · June 2026GPT-5.6 SolOpenAI · June 2026Claude Sonnet 5Anthropic · June 2026Grok 4.5xAI · July 2026GPT-5.6 TerraOpenAI · July 2026GPT-5.6 LunaOpenAI · July 2026Kimi K3Moonshot AI · July 2026Gemini 3.6 FlashGoogle DeepMind · July 2026Claude Opus 5Anthropic · July 2026Grok 4.6xAI · August 2026Gemini 3.7 FlashGoogle DeepMind · August 2026GLM-5.3Zhipu AI · August 2026Claude Fable 5.1Anthropic · September 2026Gemini 3.8 FlashGoogle DeepMind · September 2026GPT-6 AstraOpenAI · September 2026Grok 4.7xAI · September 2026GPT-6 SolOpenAI · September 2026Claude Opus 5.5Anthropic · September 2026Claude Sonnet 5.5Anthropic · September 2026GPT-6.1 SolOpenAI · September 2026n
GSM8Ksaturated92%95%96.4%96.8%92.12%92.6%6
ARC-AGI-2live89.2%92.5%90.4%90%95%93.3%6
MMLUsaturated86.4%86.8%88.7%90.4%88.6%91.8%90.8%89.5%90.1%9
AIMEcontaminated13.4%83%79.2%96.7%79.8%86.7%85.7%75.5%91.7%69.6%94.6%87%95%93%100%92.7%16
HumanEvalsaturated67%71.2%84.9%91%92%89%76.8%7
MATHsaturated42.2%60.1%76.6%71.1%73.8%94.8%97.3%98%97.4%64.5%10
MMLU-Prolive68.5%72.55%76.12%73.3%84%85.3%81.1%87.5%8
GPQAsaturating38.8%50.4%49.9%59.4%50.7%75.7%87.7%71.5%84%71.1%79.6%87.5%75.1%85.7%83.4%91.9%87%92.4%86%94.3%94.6%93.6%90.1%92%92.6%94.1%94.6%92.9%92.3%93.5%96%31
MMMUlive56.8%59.4%69.1%68.3%77.3%82.9%81.7%76.5%76.5%84.2%77.8%80.7%12
OSWorld-Verifiedsaturated14.9%61.4%66.3%76.2%75%78.7%83.4%85%85%84.8%83%11
LMArenadisputed1227 Elo1253 Elo1300 Elo1443 Elo1501 Elo1436 Elo1452 Elo1506 Elo8
SWE-benchsaturating33.2%49%48.9%71.7%49.2%63.8%72.5%65.8%74.9%69.6%77.2%76.2%80.9%80%77.8%80.6%80.6%88.6%96%19
MindCubelive48.1%51.2%57.6%59.6%64.7%59.7%58.3%62%77.3%61.2%61.7%78.3%67%60%84.1%70.4%63.9%75.6%79.9%83.9%64.9%80.1%83.1%76.8%78.1%62.6%57.8%83.3%84.3%29
OSWorld-2.0 (partial)live72.9%65.7%75.4%50.6%77.9%59%72.6%71.4%8
OSWorld-2.1 (partial)live57%74%80.7%81.8%80.1%5
ARC-AGI-1saturated9%21%82.8%15.8%24.3%66.7%65.7%63.7%80%86.2%96.5%85.7%96.5%88%87.5%98.5%16
GDPval-AA v1saturating1314 Elo1769 Elo1554 Elo1890 Elo1932 Elo1932 Elo6
SWE-bench Prolive62.1%54.2%59.1%77.8%58.6%55.4%69.2%80.3%80.3%64.6%63.2%64.7%63.4%62.7%58.7%79.2%89.9%81.3%18
MathVistasaturating49.9%50.5%63.8%67.7%71%86.8%6
ERQAlive47%53.3%60.5%60.8%50.1%58.8%50.3%59.8%70.2%54%60.7%71%54.8%56.6%74.2%64.8%58.1%61.3%70.5%75.4%59.5%66.5%67.8%71.2%63.6%63.6%62.4%66.1%70.8%29
DeepSWElive46.2%12%67%59%70%73%54%69.6%67.2%67.5%49%68.8%65.9%66.9%73.7%74.1%71%68.8%74.2%71%75.2%21
SpatialVizlive31.1%41.4%47.7%46.6%34.8%54.1%41.6%51.3%63.2%43%65.8%65.3%55.7%54.9%66.1%69.3%62.6%73.9%74.2%73.9%65.2%70.3%73.9%80.7%61.1%65.2%71.3%71.2%75.8%29
ExploitBenchlive74.2%47.9%40%78%73.5%31%52.9%33.2%70%100%80%99.7%12
IntPhys 2live53%53.1%54.7%56%54.9%56.1%54.4%52%56.9%56.3%58.3%63.4%53.6%50.8%53.6%56.4%56%60%59.9%67%55.1%66.8%67.1%63.9%61.4%61.3%58%58.6%65.1%29
HealthBenchlive43.8%51.8%56.9%60.9%66%60.5%57.8%57.7%55.7%59.8%48.5%63.4%56.7%65.6%69.2%64.2%16
GDPval-AA v2.1live1438 Elo1595 Elo1588 Elo1449 Elo1432 Elo1443 Elo1524 Elo1708 Elo1632 Elo1646 Elo1735 Elo1412 Elo1542 Elo1695 Elo1487 Elo1846 Elo1844 Elo1575 Elo18
AA Intelligence Index v4.3.2live4737.343.650.844.344.853.440.952.746.547.5585651.814
Codeforces / CodeContestslive392 Elo808 Elo1673 Elo2727 Elo2029 Elo2056 Elo3455 Elo3206 Elo8
HLE-Verifiedlive54.5%31%51.1%54.4%53.6%54.9%6
ARC-AGI-3live7.8%30.2%99.9%3
OSWorld-2.0live36.1%39.6%41.7%3
OSWorld-2.1live25.6%37.2%42.8%48.7%43.5%5
Terminal-Bench Science 0.1live24.7%22.4%29%52.6%64.6%58.7%59.9%57%8
TextQuestslive13.1%30.9%23.2%27.8%18.3%30%31%34.2%41%38.7%37%36.4%40.8%33.3%31.5%45.8%42.2%37%27.4%48.9%20.3%35.7%40.1%24.9%56.1%52.1%25.4%40.5%30.8%40%53.3%31
Terminal-Bench 4.0live37.3%12.4%23.6%51.8%20.3%11.2%55.8%19.1%57.7%37.6%66.4%70.6%12
HLE (no tools)live2.7%4.3%9.1%20.3%8.5%18.8%25.4%4.7%24.8%7.65%27.2%37.5%30.8%34.5%36.6%34.2%41.4%48.4%21.1%44.4%40.3%39%35.5%41.4%37.7%42.5%49.8%49.4%52.7%59%45.5%39.3%35.9%31.8%43.5%51%60.9%64.4%56.9%39
Terminal-Bench 3.0live34.1%34.6%15.7%26%28.3%5
Agents' Last Examlive16.4%20.5%21.1%9.2%26.6%27%25.7%30.6%27%28%30.3%28.3%27%13
AutomationBenchlive9.6%12.9%14.5%15.5%17.4%17.4%18.1%13.5%15.2%14.9%30.8%26%31.4%41.4%33.2%40%44.75%36.1%18
EnigmaEvallive0.8%5.7%11.9%5.6%7.8%10.5%6%11.7%17.8%12.4%14.5%18.3%7.6%9.1%32.4%27.6%15.6%5.5%37.2%23%20.5%16.6%41.3%39.2%17%23.1%19.1%23.6%43.9%29
covered9191912611414913441291211371113127782716103881612818722824912181714325734116135412128579