Anthropic releases Claude 2
A 100,000-token context window and a jump to 71.2% on the Codex HumanEval coding test, alongside a consumer web app opened to the US and UK.
- Models & capabilities
- Notable
Anthropic released Claude 2, its first model made available through a consumer-facing web application rather than API access alone, launching in beta in the United States and United Kingdom. The model kept the same pricing as its predecessor, Claude 1.3, on the API.
Anthropic reported gains across coding and exam benchmarks: 71.2% on the Codex HumanEval Python test, up from 56.0% for Claude 1.3, and 76.5% on the multiple-choice section of the bar exam, up from 73.0%, alongside a score above the 90th percentile on the GRE’s reading and writing sections. The model also carried a context window of up to 100,000 tokens, letting users submit inputs on the order of hundreds of pages — long enough to paste in whole documents or short books — which was, at the time, longer than what competing consumer chatbots offered by default.
On safety, Anthropic said its internal red-teaming found Claude 2 twice as likely to give a harmless response compared with Claude 1.3 on a set of harm-focused evaluations, though as with the capability figures this was a company-reported internal comparison rather than a third-party benchmark.
The release mattered less for any single number than for positioning: it was Anthropic’s first product built for a general audience rather than developers and enterprise partners, arriving roughly eight months after ChatGPT had demonstrated that a consumer chat interface, not API access, was what drove public attention to a model. Claude 2 put Anthropic directly into that market for the first time, ahead of the Claude 3 family the following spring.