An interactive radar · 58 models · 40 benchmarks · 579 sourced scores
The Race for the Frontier
Every major AI model against the benchmarks that defined the field, drawn as a radar (the same data is also on the line-chart view). Every benchmark is a spoke from the centre (easiest near the top, hardest opposite); a model reaches further out the better it scores, so each model is a polygon and the best-on-every-benchmark frontier is the outer hull. Every axis is normalised to its own 0–100%. Drag the time slider to watch the shape grow; hover a model to isolate it.
Normalised per benchmark (0–100% of each axis's own range). A model's polygon only touches the benchmarks it has a sourced score on. Hollow points are company-reported rather than independently confirmed. The full sourced-data matrix is in the table below; the same data is also on the line-chart view.
Every score, in full
The complete matrix behind the chart — 579 sourced scores across 40 benchmarks (rows, oldest and most-saturated first) and 58 models (columns, by release date), so you can see exactly what we have and where the gaps are: 25% of possible cells are filled. Italic scores are company-reported rather than independently confirmed. Hover a score for the exact setting and provenance; click it to open the source. Scroll the table both ways.