Timeline

Artificial Analysis launches independent model benchmarking site

Founded by George Cameron and Micah Hill-Smith as a side project comparing model pricing and latency, it became a widely cited independent reference.

  • Benchmarks & progress
  • Minor

Artificial Analysis began publishing continuously updated comparisons of AI models on intelligence, speed, price and other capability measures. The site was founded by George Cameron and Micah Hill-Smith, who started it as a side project in 2023 while Hill-Smith was building a legal AI research assistant and found himself repeatedly having to work out which model to use for which task on which basis of cost and latency — a comparison problem, he later said, that no independent source was tracking systematically.

Unlike benchmark releases from labs themselves, or academic leaderboards tied to a single test, Artificial Analysis aggregated multiple evaluations into composite indices and paired them with running price and throughput data pulled directly from providers’ APIs, updated as new models shipped rather than as a one-off report. The site was built and hosted as a small deployment rather than a funded venture at the outset.

It gained attention after early mentions circulated among AI-focused newsletters and podcasts in January 2024, and its release coincided with a wave of new open-weight models — including Mistral’s Mixtral 8x7B — that made third-party, apples-to-apples comparison more valuable to developers choosing between a rapidly multiplying set of options. Over the following two years it became one of the standard reference points cited by developers, journalists and, eventually, the labs themselves when discussing relative model performance — a role that also meant its methodology choices, such as which benchmarks to include or retire, drew scrutiny as the composite indices themselves evolved.