GPT-6 Astra becomes the first AI agent to win NetHack
An independent developer's harness let GPT-6 Astra ascend on its third attempt, after 37,140 turns over nearly two weeks, with wiki and source-code access and a server-verified game record.
- Models & capabilities
- Benchmarks & progress
- Notable
An agent built on OpenAI’s GPT-6 Astra won a game of NetHack — “ascending”, in the game’s terms — in what its developer, an independent programmer writing on the blog Vaguely Aligned, described as the first recorded ascension by an LLM agent. NetHack, a 1987 dungeon game with permanent death, randomised levels and thousands of interacting items, is notorious as one of the hardest games for either people or machines to finish, and had become a long-horizon test for AI agents; three days earlier, the BALROG agent leaderboard put the top configuration of Astra itself at about 13% average progress through the game — an average, the developer noted, not a win rate.
The winning run played NetHack 3.6.7 on the public Hardfought server, which keeps its own recordings and logs of every game, providing verification independent of the developer. It took 37,140 game turns spread over calendar days from 9 to 21 September, and was the third attempt: the first agent died deep in the dungeon after more than 33,000 turns, and the second at the Castle, one of the game’s late set-pieces. The agent saw the game as annotated terminal text through a custom wrapper with guards against unsafe actions, kept persistent notes on identified items, maps and its own progress, and wrote much of its own tooling along the way, including a route planner and a solver for the game’s Sokoban puzzles. The developer put the cost at about $300 in credits on top of an existing subscription.
The developer was explicit about what the result was not. The agent had unrestricted access to the NetHack wiki, spoilers and the game’s source code; a human set up the infrastructure and made suggestions; the tools and instructions changed during play; and no fixed compute budget was tracked — so it was not a closed-book benchmark evaluation. The achievement arrived three weeks after Astra’s release, as a demonstration of an agent sustaining a coherent plan across tens of thousands of steps in a hostile, open-ended environment.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 20 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseBetter Call Sol Or Better Yet Claude or Astra
- 23 September 2026 · Dylan Xu, Sebastian Prasanna, Alek Westover · Redwood ResearchAstra is much better at reasoning with filler tokens than previous models
- 24 September 2026 · Shakeel Hashim · TransformerHacking is the least worrying part of OpenAI’s Australia incident