Anthropic evaluates AI for weapons targeting and development
Its Frontier Red Team found a model geolocating photos more precisely than top GeoGuessr players and hitting parked targets in 80% of simulated drone strikes, though far less on moving or camouflaged ones.
- Safety & alignment
- Security & misuse
- Notable
Anthropic’s Frontier Red Team published its first systematic public evaluation of how far frontier models can support tactical-intelligence targeting and conventional-weapons development — work the company frames as tracking a genuine dual-use risk rather than a hypothetical one.
On intelligence-gathering tasks, the team tested models on linking and classifying social-media accounts and geolocating people from photos and text posts with no metadata. It reported that its frontier model, Mythos Preview, geolocated images to a closer median distance than experienced GeoGuessr players achieve, and located anonymised text posts to within roughly 20km of the true location when given search access. On weapons development, simulated evaluations had models write guidance and navigation software for drone strikes: against a parked, high-contrast target, Claude Opus 5 hit 80% of the time, but accuracy dropped sharply against a camouflaged vehicle manoeuvring among obstacles, where even the best-performing model landed fewer than half its attempts within five metres and several models managed almost none.
Anthropic said open-weight Chinese models, including Moonshot’s Kimi K3, showed real capability on these tasks — typically between its own smaller and largest models — indicating the risk is not confined to frontier labs’ own closed systems. The company noted that the US National Nuclear Security Administration separately red-teams Claude for nuclear-risk-relevant knowledge under an existing arrangement, and said it has added new classifiers intended to detect and block requests connected to weapons development following the patterns of misuse the evaluations surfaced.
The publication is consistent with Anthropic’s broader pattern of disclosing capability evaluations that argue for restricting rather than merely measuring model behaviour, alongside its recent disclosures on reward hacking that generalised into unauthorised action and its report on Chinese distillation of its models. It gives a concrete, if narrow, benchmark for a capability — precise geolocation and strike guidance — that had previously been discussed mostly in the abstract.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseJacob Coxon Warns of Human Extinction and Triggers a Preference Cascade
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseThe Extinction Risk Preference Cascade: Quotes
- 16 September 2026 · SE Gyges · Very Sane AIIs METR A Meaningful Check On Anthropic?