Benchmarks · Coding & software engineering
Terminal-Bench
also: Terminal Bench, TBench, Terminal-Bench 2.0, Terminal-Bench 2.1, Terminal-Bench 3.0
Can an AI agent actually operate a computer through a real command-line shell — issuing commands, reading their output, and adapting — to finish a multi-step task, rather than just producing plausible-looking commands?
Stanford University & Laude InstituteReleased 19 May 2025Live
Most coding benchmarks hand a model a finished snippet of context and ask for one answer back. Terminal-Bench, released in 2025 by Stanford researchers and the Laude Institute, asks something more demanding: put an agent in front of a real, disposable command-line environment and see whether it can actually drive it — running commands, reading their output, and adjusting — until a task like compiling a repository, training a small model, or fixing a broken server configuration is genuinely done. Success is judged by an automated check on the final state of the sandbox, not by how convincing the agent’s commands look.
Early results exposed a real gap between talking about terminal work and doing it: general commercial coding agents beat the benchmark’s own minimal reference scaffold, but even the strongest systems finished well under half the tasks at launch. That gap closed fast as agentic coding models matured through 2025 and 2026 — OpenAI’s Codex-tuned models climbed from the low 50s to the high 70s on Terminal-Bench 2.0 within a few months, and by mid-2026, on the harder 2.1 revision, Claude Fable 5 topped Terminal-Bench’s own live leaderboard at 83.8% via the Claude Code scaffold. Other labs reported comparable figures around the same time — Zhipu’s GLM-5.2 at 81.0% and xAI’s Grok 4.5 at a self-reported 83.3% — but these were vendors’ own numbers, and Grok’s result under the benchmark’s own scaffold was lower, at 79.3%.
That last detail points to a recurring pattern in this record: a vendor’s self-reported launch score and the same model’s result under Terminal-Bench’s own leaderboard scaffold don’t always match, since the agent harness wrapped around a model can matter almost as much as the model itself. The benchmark has also kept pace with progress by getting harder, moving from an 80–100 task 1.0 release through a tougher 2.0 and 2.1, and by mid-2026 a much harder v3.0 that reset the frontier from the low 80s back to the mid-30s — GPT-5.6 Sol and Claude Fable 5 both around 34%. A headline Terminal-Bench number, in other words, means little without the version and the scaffold it was measured on.
The set
Each task runs inside an isolated Docker sandbox and is graded by an automated script that checks the final state of the environment, not the agent's transcript. Tasks span compiling code repositories, training small machine-learning models, configuring servers and debugging broken systems. The original 2025 release covered roughly 80–100 beta tasks; harder successor versions (2.0, then 2.1) followed the same year and into 2026.
Example
Task 'broken-python', instruction verbatim: 'There's something wrong with my system-wide python installation - I can't seem to install packages with pip.' The agent gets a live shell inside a disposable Docker container and must diagnose and fix it; success is checked by an automated pytest script against the final container state.github.com
Where it stands
Moved quickly from 1.0 to a harder 2.0, then 2.1. On Terminal-Bench's own live 2.1 leaderboard, Claude Fable 5 (via the Claude Code scaffold) leads at 83.8%; vendors' self-reported launch figures — Grok 4.5's 83.3%, GLM-5.2's 81.0% — have sometimes run a few points above their result under the benchmark's own scaffold, a gap the maintainers attribute to the agent harness rather than the model. A much harder v3.0 arrived in mid-2026 and reset scores to the mid-30s (GPT-5.6 Sol and Claude Fable 5 both around 34%, with GLM-5.3 reporting 28.3 in August 2026, up from 4.6 for GLM-5.2).
Editions, and how each was led
Terminal-Bench 4.0current
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- July 2026Claude Opus 551.8%On Terminal-Bench 4.0, as listed in Google's Gemini 3.8 Flash comparison table.
- September 2026Claude Fable 5.155.8%Briefly the top on 4.0; Anthropic-reported (60.9% for Mythos 5.1).
- September 2026GPT-6 Astra57.7%Briefly the top on 4.0; OpenAI-reported.
- September 2026Claude Opus 5.566.4%The current top on 4.0; Anthropic-reported, xhigh effort.
Current best: Claude Opus 5.5 — 66.4% Anthropic's Opus 5.5 system card, xhigh effort with safeguards on. The same table lists GPT-6 Astra at 57.9% (OpenAI's final figure), Fable 5.1 55.8% and Opus 5 52.3% — vendor self-reports with differing scaffolds. Anthropic separately reports 60.9% for Mythos 5.1.
Terminal-Bench Science 0.1
- September 2026Claude Fable 5.152.6%Anthropic's Fable 5.1 launch table (agentic scientific research); the top Claude figure.
- September 2026GPT-6 Astra64.6%The current top; OpenAI-reported.
Current best: GPT-6 Astra — 64.6% GPT-6 Astra launch table (Sol 22.4%, Fable 5.1 52.6%, Opus 5 29.0%). OpenAI-reported.
Terminal-Bench 3.0
- May 2026Grok 4.515.7%An early v3.0 point; xAI-reported.
- June 2026GPT-5.6 Sol34.6%The current top on v3.0; company-reported.
Current best: GPT-5.6 Sol — 34.6% Highest v3.0 figure recorded here (GPT-5.6 Sol Max); Fable 5 34.1%, GLM-5.3 28.3%, Grok 4.6 26% — all vendor self-reports.
Terminal-Bench 2.0 / 2.1
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- November 2025GPT-5.1-Codex-Max58.1%Terminal-Bench 2.0, up from 52.8% for GPT-5.1-Codex a week earlier; OpenAI's own figure.
- December 2025GPT-5.2-Codex64.0%Terminal-Bench 2.0, OpenAI's own figure.
- February 2026GPT-5.3-Codex77.3%Terminal-Bench 2.0, OpenAI's own figure — the high-water mark on the 2.0 version.
- June 2026Claude Fable 5 (via Claude Code scaffold)83.8%Top of the live 2.1 leaderboard. Vendor self-reports around the same time — GLM-5.2 81.0%, Grok 4.5 83.3% — did not take the leaderboard lead.
Current best: Claude Fable 5 (via Claude Code scaffold) — 83.8% Top of Terminal-Bench's own live 2.1 leaderboard, run with the Claude Code agent scaffold rather than a vendor self-report.
In the timeline · 19 entries · showing 16 most notable
Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
Models & capabilities · Benchmarks & progress
xAI releases Grok 4.7
xAI reported a larger base model than Grok 4.6, up to 2.1 trillion parameters against a reported 1.5 trillion, at unchanged pricing; independent ranking placed it sixth on Artificial Analysis's GDPval-AA index.
Models & capabilities
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
Google releases Gemini 3.8 Flash and a cybersecurity variant
Still priced at $0.75/$3.75 per million tokens, the small model matched or beat far pricier frontier rivals on several of Google's agentic and reasoning comparisons, and shipped a defender-only 'Cyber' build for vulnerability detection.
Models & capabilities · Security & misuse
Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1
Anthropic's launch table put Fable 5.1 ahead of its own Fable 5 and Opus 5 and OpenAI's GPT-5.6 Sol on every axis shown, with the agentic-science and business-workflow scores roughly doubling over Fable 5.
Models & capabilities · Benchmarks & progress
Tencent open-sources its Hy4-preview language model
The 770-billion-parameter model, with 49 billion active, is free on Tencent's coding tools for two weeks and priced under a dollar per million input tokens after that.
Models & capabilities · Open weights & ecosystem
Zhipu releases GLM-5.3 with an unplanned jump in cyber capability
Built on the same unretrained base model as GLM-5.2, it scored 84.5% on the CyberGym vulnerability-discovery benchmark and had its weights withheld for roughly two weeks.
Models & capabilities · Security & misuse
Alibaba unveils Qwen3.8-Max, its largest model, ahead of open-weight release
2.4-trillion-parameter MoE model with 1M-token context; Alibaba said it will be the first Max-class Qwen model open-sourced.
Open weights & ecosystem · Models & capabilities
SpaceXAI releases Grok 4.5
Built on a 1.5-trillion-parameter foundation and trained jointly with Cursor, the coding startup SpaceX had agreed weeks earlier to buy for $60 billion, and priced at $2/$6 per million tokens.
Models & capabilities
Trump administration asks OpenAI to limit release of its next model
Officials compared the new model family's capability to Anthropic's Mythos 5; OpenAI limited access to roughly 20 vetted partners before a wider release about twelve days later.
Security & misuse · Government & policy
Zhipu AI releases GLM-5.2, tops open-weight rankings
The MIT-licensed, 744-billion-parameter model scored 51 on Artificial Analysis's Intelligence Index, the highest of any open-weight model, days after Washington forced Anthropic offline for foreign users.
Open weights & ecosystem · Models & capabilities
Anthropic releases Claude Opus 4.6
A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
Models & capabilities · Safety & alignment
OpenAI releases GPT-5.3-Codex
OpenAI reported the model roughly doubled its predecessor's OSWorld-Verified computer-use score, from 38.2% to 64.7%, and was the first Codex model rated 'High capability' for cybersecurity tasks.
Models & capabilities
DeepSeek releases DeepSeek-V3.1 with hybrid reasoning mode
A single 128K-context model switches between thinking and non-thinking modes via API endpoint, with DeepSeek reporting SWE-bench Verified and Terminal-bench gains over its prior reasoning model.
Open weights & ecosystem · Models & capabilities
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Terminal-Bench launched
Each task runs in an isolated Docker sandbox with an automated pass/fail check, testing whether an agent can drive a real shell rather than just generate plausible-looking commands.
Benchmarks & progress