Benchmarks · Coding & software engineering

Terminal-Bench

also: Terminal Bench, TBench, Terminal-Bench 2.0, Terminal-Bench 2.1, Terminal-Bench 3.0

Can an AI agent actually operate a computer through a real command-line shell — issuing commands, reading their output, and adapting — to finish a multi-step task, rather than just producing plausible-looking commands?

Stanford University & Laude InstituteReleased 19 May 2025Live

Most coding benchmarks hand a model a finished snippet of context and ask for one answer back. Terminal-Bench, released in 2025 by Stanford researchers and the Laude Institute, asks something more demanding: put an agent in front of a real, disposable command-line environment and see whether it can actually drive it — running commands, reading their output, and adjusting — until a task like compiling a repository, training a small model, or fixing a broken server configuration is genuinely done. Success is judged by an automated check on the final state of the sandbox, not by how convincing the agent’s commands look.

Early results exposed a real gap between talking about terminal work and doing it: general commercial coding agents beat the benchmark’s own minimal reference scaffold, but even the strongest systems finished well under half the tasks at launch. That gap closed fast as agentic coding models matured through 2025 and 2026 — OpenAI’s Codex-tuned models climbed from the low 50s to the high 70s on Terminal-Bench 2.0 within a few months, and by mid-2026, on the harder 2.1 revision, Claude Fable 5 topped Terminal-Bench’s own live leaderboard at 83.8% via the Claude Code scaffold. Other labs reported comparable figures around the same time — Zhipu’s GLM-5.2 at 81.0% and xAI’s Grok 4.5 at a self-reported 83.3% — but these were vendors’ own numbers, and Grok’s result under the benchmark’s own scaffold was lower, at 79.3%.

That last detail points to a recurring pattern in this record: a vendor’s self-reported launch score and the same model’s result under Terminal-Bench’s own leaderboard scaffold don’t always match, since the agent harness wrapped around a model can matter almost as much as the model itself. The benchmark has also kept pace with progress by getting harder, moving from an 80–100 task 1.0 release through a tougher 2.0 and 2.1, and by mid-2026 a much harder v3.0 that reset the frontier from the low 80s back to the mid-30s — GPT-5.6 Sol and Claude Fable 5 both around 34%. A headline Terminal-Bench number, in other words, means little without the version and the scaffold it was measured on.

The set

Each task runs inside an isolated Docker sandbox and is graded by an automated script that checks the final state of the environment, not the agent's transcript. Tasks span compiling code repositories, training small machine-learning models, configuring servers and debugging broken systems. The original 2025 release covered roughly 80–100 beta tasks; harder successor versions (2.0, then 2.1) followed the same year and into 2026.

Example

Task 'broken-python', instruction verbatim: 'There's something wrong with my system-wide python installation - I can't seem to install packages with pip.' The agent gets a live shell inside a disposable Docker container and must diagnose and fix it; success is checked by an automated pytest script against the final container state.github.com

Where it stands

Moved quickly from 1.0 to a harder 2.0, then 2.1. On Terminal-Bench's own live 2.1 leaderboard, Claude Fable 5 (via the Claude Code scaffold) leads at 83.8%; vendors' self-reported launch figures — Grok 4.5's 83.3%, GLM-5.2's 81.0% — have sometimes run a few points above their result under the benchmark's own scaffold, a gap the maintainers attribute to the agent harness rather than the model. A much harder v3.0 arrived in mid-2026 and reset scores to the mid-30s (GPT-5.6 Sol and Claude Fable 5 both around 34%, with GLM-5.3 reporting 28.3 in August 2026, up from 4.6 for GLM-5.2).

Editions, and how each was led

Terminal-Bench 4.0current

Released August 2026The newest, hardest edition — broader general-agent tasks; scores reset well down again. Vendor self-reports, with differing scaffolds.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. July 2026Claude Opus 551.8%On Terminal-Bench 4.0, as listed in Google's Gemini 3.8 Flash comparison table.
  2. September 2026Claude Fable 5.155.8%Briefly the top on 4.0; Anthropic-reported (60.9% for Mythos 5.1).
  3. September 2026GPT-6 Astra57.7%Briefly the top on 4.0; OpenAI-reported.
  4. September 2026Claude Opus 5.566.4%The current top on 4.0; Anthropic-reported, xhigh effort.

Current best: Claude Opus 5.5 — 66.4% Anthropic's Opus 5.5 system card, xhigh effort with safeguards on. The same table lists GPT-6 Astra at 57.9% (OpenAI's final figure), Fable 5.1 55.8% and Opus 5 52.3% — vendor self-reports with differing scaffolds. Anthropic separately reports 60.9% for Mythos 5.1.

Terminal-Bench Science 0.1

Released August 2026A science-workflow spin-off of the Terminal-Bench harness — data analysis, running simulations and fitting models from a terminal — reported by both OpenAI and Anthropic in their September 2026 launches, giving it real cross-model data.

  1. September 2026Claude Fable 5.152.6%Anthropic's Fable 5.1 launch table (agentic scientific research); the top Claude figure.
  2. September 2026GPT-6 Astra64.6%The current top; OpenAI-reported.

Current best: GPT-6 Astra — 64.6% GPT-6 Astra launch table (Sol 22.4%, Fable 5.1 52.6%, Opus 5 29.0%). OpenAI-reported.

Terminal-Bench 3.0

Released May 2026A harder v3.0 set, reported by several mid-2026 launches — numbers far below the 2.x leaderboard.

  1. May 2026Grok 4.515.7%An early v3.0 point; xAI-reported.
  2. June 2026GPT-5.6 Sol34.6%The current top on v3.0; company-reported.

Current best: GPT-5.6 Sol — 34.6% Highest v3.0 figure recorded here (GPT-5.6 Sol Max); Fable 5 34.1%, GLM-5.3 28.3%, Grok 4.6 26% — all vendor self-reports.

Terminal-Bench 2.0 / 2.1Saturating

Released November 2025The 2.0 set and its 2.1 fix — the era of the live tbench.ai leaderboard. Superseded on the frontier by the harder 3.0 and 4.0 editions, kept here for the record.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. November 2025GPT-5.1-Codex-Max58.1%Terminal-Bench 2.0, up from 52.8% for GPT-5.1-Codex a week earlier; OpenAI's own figure.
  2. December 2025GPT-5.2-Codex64.0%Terminal-Bench 2.0, OpenAI's own figure.
  3. February 2026GPT-5.3-Codex77.3%Terminal-Bench 2.0, OpenAI's own figure — the high-water mark on the 2.0 version.
  4. June 2026Claude Fable 5 (via Claude Code scaffold)83.8%Top of the live 2.1 leaderboard. Vendor self-reports around the same time — GLM-5.2 81.0%, Grok 4.5 83.3% — did not take the leaderboard lead.

Current best: Claude Fable 5 (via Claude Code scaffold) — 83.8% Top of Terminal-Bench's own live 2.1 leaderboard, run with the Claude Code agent scaffold rather than a vendor self-report.

In the timeline · 19 entries · showing 16 most notable

More coding & software engineering benchmarks