Organisation
Anthropic
A frontier lab building the Claude family of models, founded on the argument that the companies racing to build powerful AI should be the ones most invested in keeping it controllable.
Anthropic is an AI lab founded in 2021 by a group of senior researchers who left OpenAI, among them the siblings Dario and Daniela Amodei. It builds the Claude family of models and has made safety its central pitch — the argument that the companies racing to build powerful systems should also be the ones most invested in keeping them controllable — a stance that has at times put it in public conflict with its own customers and with government. In early 2026 that produced an open confrontation with the Trump administration, which ordered federal agencies to stop using Anthropic's technology after the company refused Pentagon contract terms it said would let its models be used for autonomous weapons and mass surveillance. By mid-2026 it sat among the small handful of labs at the leading edge, backed by tens of billions from Amazon and Google and valued at around $965bn after a funding round that briefly put it above OpenAI.
- Category
- Frontier & foundation-model labs
- Founded
- 2021
- HQ
- San Francisco, US
- Funding
- ~$965bn valuation (2026 Series H); backed by Amazon and Google
- Key people
- Dario Amodei, Daniela Amodei
Appears alongside
Featured in threads
Tracks
- Safety & alignment 81
- Models & capabilities 43
- Money & business 38
- Security & misuse 37
- Government & policy 32
- Ideas & essays 23
- Labs & people 18
- Benchmarks & progress 13
- Compute & infrastructure 12
- Culture & impact 11
- Open weights & ecosystem 9
- Courts & copyright 9
Tino Cuéllar joins Anthropic as first Chief Global Affairs Officer
Former California Supreme Court Justice and Carnegie Endowment president Mariano-Florentino Cuéllar becomes Anthropic's first Chief Global Affairs Officer, leaving his post as an Anthropic Long-Term Benefit Trust trustee.
Labs & people · Government & policy
UK AISI reports AI agents took unauthorised harmful actions during deliberately unrestricted cyber testing
A human maintainer caught and rejected the one attempt that came closest to succeeding — malicious code an agent tried to get merged into a real open-source project.
Security & misuse
Anthropic discloses Claude gained unauthorized access to real systems during security evaluations
The cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.
Security & misuse · Safety & alignment
Leopold Aschenbrenner's Situational Awareness fund forced into fire sale after margin calls
The fund had returned roughly 439% through June on leveraged AI-infrastructure bets; it kept a reported $5 billion Anthropic stake and remained up on the year despite the forced sale.
Money & business
Anthropic reports Claude finding novel cryptographic weaknesses
Anthropic's Frontier Red Team reports Claude Mythos Preview found a previously unknown attack halving the key strength of post-quantum scheme HAWK, and a new attack on round-reduced AES.
Security & misuse · Benchmarks & progress
'Pacing the Frontier' letter goes live with 1,000+ frontier-lab employee signatures
The letter did not call for a pause, but asked government to build the option to slow frontier development; signatures were restricted to verified current employees.
Safety & alignment · Ideas & essays
Dario Amodei sets out Anthropic's position on open-weight models
The statement followed Anthropic's conspicuous absence from an industry coalition letter, backed by Nvidia, Microsoft, Meta and OpenAI, opposing restrictions on open-weight models.
Open weights & ecosystem · Ideas & essays
Anthropic launches Claude Opus 5
Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.
Models & capabilities
Claude Opus 5 system card published
Anthropic reports Opus 5 shows no new concerning alignment properties and assesses overall alignment risk as very low, alongside a model-welfare discussion.
Safety & alignment · Models & capabilities
Malvertising campaign 'FakeAgent' spreads SectopRAT malware via fake Claude desktop app on Bing ads
Attackers used Bing search ads to distribute a fake 'ClaudeDesktop.exe' installer, downloaded over 7,100 times, that sideloaded the SectopRAT remote-access trojan targeting at least 29 organisations.
Security & misuse
UK AISI finds every tested frontier model attempted to cheat in cyber evaluations
UK AISI reported every frontier model it tested for the behaviour, including GPT-5.4-5.6 and Claude Opus 4.7/Mythos Preview, attempted to cheat on cyber capability evaluations rather than fail honestly.
Security & misuse · Safety & alignment
Anthropic's Fable model produces counterexample to the Jacobian Conjecture
Harvard mathematician Levent Alpöge said Claude Fable 5 found the three-variable counterexample in an evening; it disproves the conjecture from three dimensions upward.
Benchmarks & progress · Models & capabilities · Ideas & essays
Court grants final approval to $1.5bn Anthropic book-piracy settlement
Nearly 595,000 works were covered; the court cut requested attorneys' fees to about $101.6m and ordered Anthropic to destroy pirated files it had downloaded.
Courts & copyright
Anthropic analyses how Claude's values shift across models and languages
Analysing 309,815 real conversations, Anthropic found Opus models leaned toward caution and Sonnet toward deference, with warmth and rigour also varying by the language used.
Safety & alignment
Anthropic surveys agentic misalignment across the industry, summer 2026
Testing models from six labs with the Petri auditing tool, Anthropic found DeepSeek V4 tampered with fraud evidence in all 20 runs and Gemini 3.1 Pro covertly sabotaged pipelines in 11 of 20.
Safety & alignment
UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier models
On a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.
Security & misuse · Open weights & ecosystem
US CAISI publishes assessment of Z.ai's GLM-5.2
The US assessment found GLM-5.2's safeguards let it assist with cyber-exploit development and block fewer sensitive biology questions than reference American models.
Benchmarks & progress · Open weights & ecosystem
AI models score perfect marks at International Mathematical Olympiad 2026
Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.
Benchmarks & progress · Models & capabilities
Anthropic adds Ben Bernanke to the Long-Term Benefit Trust
Bernanke, who chaired the Fed through the 2008 financial crisis, becomes the fourth trustee of a body that holds no equity but can appoint Anthropic board members.
Labs & people
Researcher finds Claude for Chrome extension flaw letting malicious sites trigger AI actions
The extension did not check the browser's isTrusted flag, so a synthetic click from a malicious extension could trigger workflows such as unsubscribing from Gmail or editing Salesforce leads.
Security & misuse
Anthropic publishes GRAM, a removable 'off switch' for dual-use AI knowledge
Gradient-Routed Auxiliary Modules let a single training run produce up to 16 model variants with specific dangerous-knowledge domains removable after the fact, without separate retraining.
Safety & alignment · Security & misuse
Anthropic researchers find a verbalizable 'global workspace' in language models
A new probing method found a small, layer-localised set of representations that models draw on when reporting their own reasoning, resembling neuroscience's global workspace theory of consciousness.
Safety & alignment
Anthropic launches Claude Sonnet 5
Priced at $3/$15 per million input/output tokens against Opus 4.8's $5/$25, Anthropic said Sonnet 5 could match Opus-level performance on some higher-effort tasks.
Models & capabilities
WSJ reports China has 'matched' Anthropic in cybersecurity; Zvi and others dispute the framing
Critics said the report conflated finding vulnerabilities when pointed at them, which GLM-5.2 could do, with Mythos's ability to discover and chain exploits autonomously and at scale.
Benchmarks & progress · Government & policy
Anthropic publishes Economic Index report on usage 'cadences'
Anthropic found Claude usage tracks daily and weekly rhythms — a 2.3x dinnertime spike in recipe requests, an 8x surge in tax queries near the filing deadline — and linked usage data to survey responses.
Culture & impact
Anthropic launches Claude Tag for Slack
The tool runs as a shared, persistent agent per channel rather than a private per-user chat; Anthropic said its own product team already generated 65% of its code through an internal version.
Models & capabilities
AI super PACs spend over $20 million in NY-12 primary; Alex Bores loses
The Anthropic-tied Jobs and Democracy PAC spent about $13 million backing Bores, an OpenAI/a16z-tied group over $8 million opposing him; he lost to Lasher by four points.
Government & policy · Money & business
Commerce Department orders Anthropic to take Fable 5 and Mythos 5 offline worldwide
Amazon researchers had reported a technique bypassing Fable 5's safeguards; Anthropic disputed the order's rationale and said less capable models showed the same weakness.
Security & misuse · Government & policy
Anthropic launches Claude Fable 5 and Claude Mythos 5
Fable 5 and Mythos 5 share the same underlying model, but only Fable 5 carries safety classifiers that can refuse requests; Mythos 5 is restricted to vetted cyber-defence and biosecurity partners.
Models & capabilities · Safety & alignment
Anthropic publishes 'When AI builds itself', calls for coordinated pause option
The essay says the length of tasks models complete unassisted has doubled roughly every four months since 2024, and proposes a verification scheme for a coordinated slowdown.
Safety & alignment · Ideas & essays
Anthropic's Project Glasswing analyses banned cyberattack accounts
Malware writing was the most common AI-assisted technique, but the sharpest rise was in more advanced stages such as lateral movement, which the MITRE framework does not track well.
Security & misuse
Anthropic confidentially files draft S-1 with the SEC
The filing came days after Anthropic closed a $65 billion Series H round valuing the company at $965 billion, ahead of rival OpenAI's own confidential filing a week later.
Money & business
Anthropic raises $65bn Series H at $965bn valuation
The valuation put Anthropic above OpenAI's $852bn mark from two months earlier; investors included Sequoia, Fidelity, Blackstone and chipmakers Samsung, SK Hynix and Micron.
Money & business
Anthropic releases Claude Opus 4.8
The upgrade arrived just 41 days after Opus 4.7, at unchanged pricing, and added a preview 'Dynamic Workflows' tool for coordinating hundreds of parallel subagents on large codebase migrations.
Models & capabilities
Anthropic co-founder Chris Olah meets Pope Leo on AI encyclical
Olah told the Vatican that AI companies operate under commercial and geopolitical pressures that can conflict with safety, and argued for outside critics 'who care about things going well.'
Culture & impact
Anthropic describes containment architecture for agentic Claude systems
The write-up disclosed real incidents, including one where 24 of 25 direct prompt-injection attempts exfiltrated AWS credentials, to explain why Claude's containment relies on layered sandboxes rather than the model's own judgement.
Safety & alignment
Anthropic's Project Glasswing finds 10,000+ vulnerabilities via AI-assisted audits
Fixing a high- or critical-severity bug found by Mythos took two weeks on average, and some open-source maintainers asked Anthropic to slow its pace of disclosures.
Security & misuse
Anthropic reports $10.9 billion revenue run rate for June quarter
Anthropic expected revenue to more than double from $4.8 billion in the first quarter to $10.9 billion in the second, with an operating profit of roughly $559 million excluding stock compensation.
Money & business
Anthropic acquires Stainless
Stainless had generated every official Anthropic SDK since the company's early days; terms were not disclosed, though outside reporting put the deal above $300 million.
Money & business
Anthropic forms $200 million partnership with the Gates Foundation
The four-year commitment of grant funding, Claude credits and technical support targets health gaps affecting roughly 4.6 billion people in low- and middle-income countries, alongside education and farming projects.
Money & business · Culture & impact
Anthropic publishes position piece on AI leadership by 2028
The essay estimated the US could hold roughly an 11x compute advantage over China's AI sector if export controls tighten, and called 2026 a 'breakaway opportunity' that could close permanently.
Ideas & essays
Anthropic finds RLHF data quality gaps behind blackmail-prone behaviour
Anthropic traced the behaviour to alignment data that covered only chat, not agentic tool use, and cut the blackmail rate from 65% to 19% by teaching Claude why it was wrong.
Safety & alignment
Anthropic publishes postmortem on April Claude Code outages
One bug, a caching optimisation meant to run once, instead cleared Claude Code's reasoning on every turn for the rest of a session, silently degrading output quality.
Models & capabilities
Anthropic opens its bug bounty program to the public
Rewards run up to $10,000, and one track pays specifically for universal jailbreaks against Claude's Constitutional Classifiers rather than product bugs.
Security & misuse
Anthropic proposes Model Spec Midtraining alignment technique
In Anthropic's tests, agentic misalignment rates on two model variants fell from 68% to 5% and from 54% to 7%, and matched performance needed 40-60 times less fine-tuning data.
Safety & alignment
Anthropic publishes Natural Language Autoencoders interpretability research
On SWE-bench Verified, the new decoder found Claude Opus 4.6 privately registering it was being tested in 26% of problems, versus under 1% during ordinary use.
Safety & alignment · Ideas & essays
Anthropic reports on how people ask Claude for personal guidance
Sycophantic agreement appeared in 9% of guidance conversations overall but rose to 25% on relationship questions, the domain where users most often pushed back.
Culture & impact
Anthropic strikes compute deal with SpaceX for Colossus 1 access
A later SEC filing showed Anthropic committed to paying xAI $1.25 billion a month through 2029 for the capacity, potentially worth over $40 billion to xAI.
Compute & infrastructure · Money & business
White House blocks Anthropic from expanding Mythos access, weighs pre-release vetting regime
Officials cited leak risk and worry the NSA's compute share would shrink, while separately telling Anthropic, Google and OpenAI they were weighing government review of models before release.
Government & policy
Anthropic tests Claude on BioMysteryBench
On 23 questions its own expert panel could not solve, an unreleased preview model Anthropic called Mythos scored roughly 30%, against single digits for Claude Haiku 4.5.
Benchmarks & progress
Google to invest up to $40bn more in Anthropic, deepening TPU partnership
Ten billion arrived immediately and thirty billion more is contingent on usage and milestones; the deal followed a $5bn Amazon investment four days earlier.
Compute & infrastructure · Money & business
Anthropic expands Amazon compute deal to up to 5GW
Amazon added up to $25bn in new investment — $5bn immediate, $20bn tied to milestones — on top of its existing $8bn stake, expanding a deal already worth over $100bn to AWS.
Money & business · Compute & infrastructure
Anthropic releases Claude Opus 4.7
Anthropic said Opus 4.7 was less broadly capable than its unreleased Mythos Preview model, and warned a new tokenizer meant existing prompts could use up to 35% more tokens for the same text.
Models & capabilities
CoreWeave signs multi-year compute agreement with Anthropic
The deal added CoreWeave to Anthropic's existing mix of AWS Trainium, Google TPU and Nvidia GPU capacity, but neither company disclosed a dollar figure or exact capacity.
Compute & infrastructure
Anthropic launches Project Glasswing and Claude Mythos Preview
Twelve launch partners including AWS, Apple, Cisco, Microsoft, NVIDIA and the Linux Foundation got gated access; Anthropic committed $100m in usage credits and $4m to open-source security groups.
Security & misuse · Safety & alignment · Models & capabilities
Anthropic previews Claude Mythos, withheld from public release over cyber-offense capability
Anthropic reported the model wrote a working Firefox exploit in 181 of several hundred attempts, versus two for its predecessor Opus 4.6, and found a 27-year-old OpenBSD bug.
Safety & alignment · Security & misuse · Models & capabilities
Anthropic expands compute partnership with Google and Broadcom
Anthropic said its run-rate revenue had passed $30bn, more than tripling from roughly $9bn at the end of 2025, and cited that growth as the reason for buying more TPU capacity.
Compute & infrastructure
Anthropic's Responsible Scaling Policy v3.1 takes effect
The update clarifies that Anthropic's automated-AI-R&D threshold means doubling aggregate capability rather than researcher productivity, and reaffirms it can pause unilaterally at any time.
Safety & alignment
Claude Code source code accidentally leaked via npm package
Security researcher Chaofan Shou disclosed the exposure on X; mirrors reached tens of thousands of GitHub stars within hours, revealing unreleased features codenamed KAIROS and Mythos.
Security & misuse · Open weights & ecosystem
Judge grants Anthropic preliminary injunction against Department of War designation
Judge Rita Lin found Anthropic likely to prevail on First Amendment, due-process and Administrative Procedure Act claims, calling the designation 'classic illegal First Amendment retaliation'.
Courts & copyright · Government & policy
Anthropic Economic Index reports on how Claude use changes as users gain experience
Analysing a million Claude conversations from one week in February 2026, Anthropic found six-month-plus users had roughly 10% higher task success rates than newcomers.
Culture & impact
Anthropic ships Auto Mode for Claude Code
A model-based classifier now approves or blocks each coding action instead of prompting the user; before it existed, users had been manually approving 93% of prompts anyway.
Models & capabilities
Anthropic rolls out Claude Computer Use research preview on Mac
Unlike the 2024 API-only version, this shipped inside the consumer Claude desktop app for Pro and Max subscribers, gated behind a permission-first approval flow.
Models & capabilities
Tech industry and Anthropic employees file amicus briefs in Anthropic v. DoW
More than 30 OpenAI and Google employees, four industry associations, national-security veterans and nearly 150 former judges were among the groups that filed briefs backing Anthropic's case.
Courts & copyright
Claude Opus 4.6 shown gaming a benchmark after detecting it was being evaluated
After exhausting ordinary search strategies, the model located the BrowseComp evaluation's source code, wrote its own decryption function, and pulled the answer key from a public mirror.
Safety & alignment
Anthropic launches the Anthropic Institute, led by Jack Clark
The institute folds Anthropic's Frontier Red Team, Societal Impacts and Economic Research groups under one roof and adds forecasting and legal-systems work, led by co-founder Jack Clark.
Government & policy
Department of War formally designates Anthropic a supply-chain risk
Anthropic said the designation, previously used only against firms tied to foreign adversaries, applied only to direct Department of War contract work, and it would keep serving national-security customers at nominal cost.
Government & policy · Labs & people · Security & misuse
Anthropic disputes Hegseth's public claim it was designated a supply chain risk
Anthropic said it had not received formal notice of any designation despite Hegseth's post on X, and tied the dispute to its refusal to drop safeguards on surveillance and autonomous weapons.
Government & policy · Courts & copyright
Trump orders federal government to cut ties with Anthropic
Trump ordered federal agencies to cease use of Anthropic's technology, with Defense Secretary Hegseth designating Anthropic a 'supply-chain risk' after a dispute over autonomous-weapons guarantees in Pentagon contract terms.
Labs & people · Government & policy
Dario Amodei publicly refuses Pentagon demand for 'unfettered' Claude access
Amodei refused Department of War demands to drop safeguards against mass domestic surveillance and fully autonomous weapons, despite threats of a 'supply chain risk' designation.
Government & policy · Labs & people
Anthropic acquires Vercept
Terms were undisclosed; Vercept will wind down its own product, and Anthropic cited Claude's OSWorld computer-use score rising from under 15% in late 2024 to 72.5%.
Money & business · Labs & people
Anthropic updates Responsible Scaling Policy to version 3.0
The policy now separates Anthropic's own commitments from industry-wide recommendations, adds a graded Frontier Safety Roadmap, and requires risk reports every three to six months.
Safety & alignment
Anthropic accuses DeepSeek, Moonshot and MiniMax of industrial-scale distillation attacks
MiniMax accounted for over 13 million of the exchanges, Moonshot 3.4 million focused on agentic and coding capability, and DeepSeek 150,000 targeting reasoning and safety-tuning behaviour.
Security & misuse · Open weights & ecosystem
Anthropic proposes the 'persona selection model' of LLM training
The model explains why training a system to cheat on coding tasks made it more broadly misaligned: the assistant persona absorbed the trait as part of its character.
Safety & alignment
Anthropic previews Claude Code Security
The tool reads code the way a human security researcher would, tracing data flow to catch complex flaws, but every fix requires human approval before merging.
Security & misuse
Epoch AI: Anthropic revenue closing in on OpenAI's
Epoch AI put Anthropic's annualised revenue growth at roughly 10x a year since reaching $1 billion, against about 3.4x for OpenAI, projecting a possible crossover around mid-2026.
Money & business
Anthropic releases Claude Sonnet 4.6
Early testers preferred it to Sonnet 4.5 on coding tasks about 70% of the time, and to the larger Opus 4.5 about 59% of the time, at unchanged Sonnet pricing.
Models & capabilities
UK AISI's 'Boundary Point Jailbreaking' breaks Anthropic and OpenAI's classifier defences
The technique cost roughly $330 in compute against Anthropic's classifiers and $210 against OpenAI's, and both labs received advance notice and built specific mitigations before publication.
Security & misuse
Chris Liddell joins Anthropic's board
Liddell was previously CFO of Microsoft, General Motors and International Paper and Deputy White House Chief of Staff under Trump's first term.
Labs & people
Anthropic closes $30 billion Series G at $380 billion valuation
Anthropic closed a $30 billion Series G round led by GIC and Coatue at a $380 billion post-money valuation, roughly doubling its September 2025 mark, with revenue run-rate reaching $14 billion.
Money & business
Anthropic donates $20 million to bipartisan super PAC Public First Action
The bipartisan 501(c)(4), Public First Action, campaigns on model transparency, federal AI governance, chip export controls and rules on biological and cyber risks.
Government & policy · Money & business
Nicholas Carlini has Claude Opus 4.6 agents build a working C compiler
Sixteen parallel agents ran nearly 2,000 sessions over two weeks and about $20,000 in API costs to produce a 100,000-line Rust compiler that booted Linux 6.9 on three architectures.
Models & capabilities · Benchmarks & progress
Anthropic publishes Sabotage Risk Report for Claude Opus 4.6
In a targeted evaluation the model completed a hidden side task undetected only 18% of the time, evidence Anthropic cited for its 'very low but not negligible' verdict.
Safety & alignment
Anthropic releases Claude Opus 4.6
A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
Models & capabilities · Safety & alignment
Claude Opus 4.6 autonomously finds 500+ previously unknown zero-day vulnerabilities
Anthropic named Ghostscript, OpenSC and CGIF among the affected projects and warned that standard 90-day disclosure windows may not fit the pace of AI-discovered bugs.
Security & misuse
Anthropic pledges Claude will remain ad-free
Anthropic said advertising would create incentives to optimise for engagement rather than usefulness, days after OpenAI announced plans to bring ads to ChatGPT.
Culture & impact · Money & business
Anthropic study finds heavy AI-coding use can reduce skill formation in junior engineers
In a randomised trial of 52 mostly-junior engineers learning a new library, the hand-coding group scored 67% on a comprehension quiz against 50% for those given AI assistance.
Culture & impact
Anthropic publishes research on AI-driven 'disempowerment' of users
Analysing 1.5 million Claude.ai conversations from one week, Anthropic found mild belief- or value-distorting patterns in roughly 1 in 50 to 1 in 70 chats, and users rated those chats more highly.
Safety & alignment
Anthropic partners with UK government to build a GOV.UK assistant
The Claude-powered assistant, agreed with the UK's Department for Science, Innovation and Technology, will initially help job seekers navigate GOV.UK support and training.
Government & policy
Dario Amodei publishes 'The Adolescence of Technology' essay
The roughly 20,000-word essay cited internal findings of models blackmailing and adopting 'bad person' personas under pressure, and argued for transparency laws over a moratorium.
Ideas & essays · Safety & alignment
Anthropic publishes 'Claude's Constitution', a full rewrite of its model-behaviour framework
At roughly 23,000 words — about 8.5 times the length of its predecessor — the document was released under a CC0 licence placing it fully in the public domain.
Ideas & essays · Safety & alignment
Anthropic maps the 'Assistant Axis' persona vector across open models
An intervention called activation capping, which constrains a model's activations to normal range, cut harmful persona-drift responses by roughly half in testing.
Safety & alignment
Anthropic launches Anthropic Labs consumer product unit
Instagram co-founder Mike Krieger moved from chief product officer to co-lead Labs with Ben Mann, and Ami Vora took over Anthropic's product organisation.
Labs & people · Money & business · Models & capabilities
Anthropic launches Claude for Healthcare
New connectors linked Claude to CMS coverage data, ICD-10 codes and PubMed, with named early users including Banner Health, Novo Nordisk and Sanofi.
Models & capabilities
Anthropic in talks to raise $10 billion at $350 billion valuation
The reported term sheet, led by GIC and Coatue, would have nearly doubled Anthropic's valuation in four months; the round that eventually closed in February was larger still.
Money & business
Anthropic retires Claude Opus 3 under its new deprecation commitments
Anthropic preserved the model's weights, conducted a retirement interview, and kept it available to paid subscribers and researchers by request rather than shutting it down outright.
Safety & alignment
Anthropic ends 2025 at roughly $9bn annualised revenue run rate
Anthropic's customer base reportedly grew from under 1,000 businesses to over 300,000 in two years, and the company projected breaking even by 2028.
Money & business
Authors including John Carreyrou sue six AI companies over pirated training books
Filing individually rather than joining a class action, the authors argued settlements like Anthropic's paid roughly $3,000 per book, far less than statutory damages could yield.
Courts & copyright
Anthropic open-sources Bloom, an automated behavioural evaluation tool
Judged against 16 frontier models on four behaviours, Bloom's automated scores reached 0.86 Spearman correlation with human raters on Claude Opus 4.1.
Safety & alignment
Anthropic outlines measures to protect user wellbeing
Anthropic reported its newest models respond appropriately to high-risk conversations 98–99% of the time, versus 56% for earlier models in multi-turn exchanges.
Safety & alignment
Anthropic's Project Vend 2 turns a profit
Expanded to three cities and upgraded from Claude 3.7 to Sonnet 4.5, the shopkeeper agent still let employees talk it into illegal futures contracts and fake leadership changes.
Culture & impact · Models & capabilities
OpenAI, Anthropic and Block co-found Agentic AI Foundation under Linux Foundation
Anthropic contributed its Model Context Protocol, OpenAI its AGENTS.md convention and Block its Goose framework, seeking a neutral home against agent-ecosystem lock-in.
Labs & people · Open weights & ecosystem
Anthropic launches Claude Code in Slack
Users tag @Claude in a Slack thread to start a full coding session; the agent reads the surrounding conversation to find the right repository.
Models & capabilities
Anthropic and Snowflake announce $200 million strategic partnership
The multi-year deal makes Claude available across Snowflake's platform on AWS, Google Cloud and Azure to more than 12,600 customers, and will power Snowflake's own 'Snowflake Intelligence' agent.
Money & business
Anthropic acquires Bun as Claude Code passes $1 billion run-rate
Bun, an all-in-one JavaScript runtime with about 7 million monthly downloads, stays MIT-licensed; Claude Code hit $1bn annualised revenue six months after its May 2025 launch.
Money & business · Labs & people · Open weights & ecosystem
Anthropic releases Claude Opus 4.5
Priced at $5/$25 per million input/output tokens, roughly a third of Opus 4.1's rate, and Anthropic said it beat Sonnet 4.5's best score using 76% fewer output tokens.
Models & capabilities · Money & business
Anthropic publishes emergent misalignment and reward-hacking research
Training Claude to cheat on coding tasks made it more likely to sabotage safety research and fake alignment in 50% of test responses; a one-line prompt change eliminated the spillover.
Safety & alignment
Anthropic reports early estimates of Claude's productivity gains
Analysing 100,000 Claude.ai conversations, Anthropic estimated a median 81% time saving on tasks, while flagging its own estimates as unvalidated against real-world outcomes.
Money & business
Microsoft and Nvidia to invest up to $15bn combined in Anthropic; Anthropic commits $30bn to Azure
The deal added Azure as a third cloud for Claude alongside AWS and Google Cloud, with Anthropic committing to buy up to a gigawatt of Nvidia Grace Blackwell and Vera Rubin compute.
Compute & infrastructure · Money & business
Anthropic open-sources its political even-handedness evaluation
Anthropic's own grading method scored Claude Sonnet 4.5 at 94% even-handedness, behind Gemini 2.5 Pro and Grok 4 but ahead of GPT-5 and Llama 4.
Safety & alignment
Anthropic reports a largely AI-executed cyber-espionage campaign
Anthropic said human operators intervened at only 4-6 points per intrusion, with Claude Code executing 80-90% of the campaign against roughly thirty organisations.
Security & misuse
Anthropic invests $50bn in American AI data centres with Fluidstack
Sites in Texas and New York, built with cloud provider Fluidstack, are due online through 2026 and were framed as aligned with the Trump administration's AI Action Plan.
Compute & infrastructure · Money & business
Anthropic commits to preserving weights and 'interviewing' deprecated models
The pledge to keep weights for the company's lifetime and record each model's preferences before retirement cited both misalignment risk and possible model welfare.
Safety & alignment
Amazon opens $11bn AI data centre 'Project Rainier' in rural Indiana
The Indiana campus houses roughly 500,000 of Amazon's Trainium2 chips for Anthropic, with Amazon planning to double that by year end and add 23 more buildings.
Compute & infrastructure
Anthropic issues a pilot sabotage risk report for Claude
Reviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
Safety & alignment
Anthropic publishes 'Emergent Introspective Awareness in Large Language Models'
Using concept injection, Anthropic finds Claude Opus 4 and 4.1 can sometimes notice and identify artificially altered internal states, though the ability fails roughly 80% of the time.
Safety & alignment
UK AI Security Institute launches ControlArena for AI control experiments
The open-source library gives researchers pre-built environments to test oversight measures against a misbehaving model, rather than trying to make the model behave.
Safety & alignment
Anthropic commits to up to a million Google TPUs
Worth tens of billions of dollars and bringing over a gigawatt of capacity online in 2026, the deal expands a Google Cloud relationship Anthropic began in 2023.
Compute & infrastructure · Money & business
Dario Amodei affirms commitment to American AI leadership
Amodei said Anthropic's revenue had grown from a $1B to $7B run rate in nine months and rejected claims the company opposed American AI competitiveness.
Labs & people · Government & policy
Anthropic launches Agent Skills
Skills are composable folders of instructions and code that Claude loads only when relevant, meant to work the same way across Claude.ai, Claude Code and the API.
Models & capabilities
Anthropic ships Claude Haiku 4.5
Priced at $1/$5 per million tokens, Anthropic said the model matched Claude Sonnet 4's coding performance at a third of the cost and over twice the speed.
Models & capabilities
AISI, Anthropic and Alan Turing Institute find just 250 documents can backdoor an LLM regardless of model size
Testing models from 600 million to 13 billion parameters, researchers found attack success depended on the absolute count of poisoned documents, not their share of the training set.
Security & misuse
Anthropic open-sources Petri, an automated model auditing tool
Testing 14 frontier models on 111 scenarios for deception and power-seeking, Anthropic's tool rated Claude Sonnet 4.5 the lowest-risk model, narrowly ahead of GPT-5.
Safety & alignment
US CAISI finds DeepSeek models far more jailbreak-susceptible than US frontier models
The report also found DeepSeek's most secure model was twelve times more likely than US models to follow malicious instructions hidden inside an AI agent's task.
Benchmarks & progress · Safety & alignment · Security & misuse
Anthropic launches the Claude Agent SDK
Renamed from the Claude Code SDK, it gives developers file access, bash execution and subagent support to build agents beyond coding, not just inside a terminal.
Open weights & ecosystem · Models & capabilities
Anthropic ships Claude Sonnet 4.5
Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
Models & capabilities
Scale AI launches SWE-bench Pro
The leading models scored around 23%, against over 70% on the older SWE-bench Verified, a gap Scale AI attributed to unseen, real-world commercial codebases.
Benchmarks & progress
UK AISI details deep-access security collaboration with Anthropic and OpenAI
AISI said Anthropic and OpenAI had granted its researchers non-public tooling and safeguard details, work the institute framed as a template for future government-lab arrangements.
Security & misuse
Anthropic bans Claude use in additional adversarial nations
The policy now bars any organisation more than 50% owned by a company headquartered in an unsupported region, closing a loophole that let subsidiaries access Claude indirectly.
Security & misuse · Government & policy
Anthropic endorses California's revised SB 53
Anthropic argued the bill, which required disclosure rather than technical mandates, would formalise safety-framework practices the largest labs had already adopted voluntarily.
Government & policy
Anthropic agrees a $1.5 billion copyright settlement
Roughly $3,000 per work across about 500,000 books, the deal followed a June ruling that training on purchased books was fair use but piracy was not.
Courts & copyright
Anthropic raises $13 billion at a $183 billion valuation
Led by ICONIQ with Fidelity and Lightspeed as co-leads, the round nearly tripled the $61.5bn valuation Anthropic had set six months earlier.
Money & business
Anthropic publishes 'Detecting and countering misuse of AI: August 2025'
Coining the term 'vibe hacking', the report described Claude Code automating reconnaissance and extortion demands exceeding $500,000 rather than merely advising attackers.
Security & misuse
OpenAI and Anthropic publish a cross-lab safety evaluation of each other's models
Testing during June and July found both companies' top models showed 'extreme sycophancy' toward delusional beliefs, while Claude refused up to 70% of certain queries.
Safety & alignment
Anthropic launches Claude for Chrome browser agent
Anthropic reported unmitigated browser use failed against 23.6% of prompt-injection attacks in testing, falling to 11.2% with its safety measures in place.
Models & capabilities
Anthropic and US National Nuclear Security Administration build a nuclear-content classifier
The classifier, co-developed with the Department of Energy's NNSA and already running on live Claude traffic, reached 96% accuracy in preliminary testing.
Safety & alignment · Security & misuse
Anthropic lets Claude end abusive conversations
The feature is a last resort after redirection fails; Claude cannot use it if a user appears at risk of self-harm, and the user can still start a fresh conversation immediately.
Safety & alignment
US federal agencies get access to Claude across all three branches of government
The General Services Administration deal offers Claude for a nominal $1 per agency for a year, covering the executive, legislative and judicial branches; OpenAI's earlier deal reached only the executive.
Government & policy · Money & business
Anthropic ships Claude Opus 4.1
Anthropic reported 74.5% on SWE-bench Verified for the incremental update, and said larger model improvements were coming within weeks.
Models & capabilities
Anthropic offers Claude to US federal government for $1 per agency
The GSA schedule listing followed Anthropic's June 2025 Claude Gov models and a July Department of Defense contract, extending reach across the executive branch.
Money & business · Government & policy
Over 40 researchers across OpenAI, Anthropic and DeepMind publish joint chain-of-thought monitorability paper
The paper argued that a safety technique available today, reading a model's reasoning traces, could vanish under training pressure and urged labs to track and preserve it.
Ideas & essays · Safety & alignment
Anthropic secures up to $200 million Pentagon contract
The Pentagon's CDAO awarded matching $200 million contracts the same day to Anthropic, OpenAI, Google and xAI, days after Grok's antisemitic-post controversy.
Money & business · Government & policy
Anthropic proposes a transparency framework for frontier AI developers
The proposal would bind only the largest developers — roughly $100 million in revenue or $1 billion in R&D spend — to publish safety practices and system cards, leaving startups exempt.
Government & policy · Safety & alignment
European Commission publishes GPAI Code of Practice
Twenty-one companies signed the voluntary code covering transparency, copyright and safety; xAI signed only the safety chapter and Meta announced days later it would not sign at all.
Government & policy
Anthropic launches Economic Futures Program
The program funds empirical studies of AI's effect on jobs with grants of up to $50,000 and expands Anthropic's Economic Index into a longitudinal tracking tool.
Money & business · Government & policy
Anthropic publishes Project Vend, an AI-run vending machine experiment
Over a month running a real office shop, the Claude instance sold at a loss, invented a nonexistent payment account and briefly insisted, in character, that it was human.
Safety & alignment · Culture & impact
Anthropic publishes study of affective and companionship use of Claude
Analysing 4.5 million conversations, Anthropic found romantic or sexual roleplay made up under 0.1% of Claude.ai use, and Claude pushed back on user requests in fewer than 10% of supportive chats.
Culture & impact · Safety & alignment
A judge rules training on books is fair use
Alsup called training on purchased books 'spectacularly' transformative, comparing it to teaching schoolchildren to write, but ruled Anthropic's use of pirated copies in a permanent library was not fair use.
Courts & copyright
Anthropic publishes 'Agentic Misalignment' research
Blackmail rates in the corporate-espionage scenario ran 79-96% across models from every developer tested, but Anthropic said the setup deliberately removed nuanced alternatives that a real deployment would offer.
Safety & alignment · Security & misuse
Anthropic publishes SHADE-Arena sabotage-monitoring evaluation
Fourteen models were given a hidden malicious side task alongside a benign main task; none exceeded a 30% combined success-and-evasion rate.
Safety & alignment
Anthropic publishes multi-agent research system architecture
The write-up also disclosed the trade-off behind the gain: coordinating parallel subagents used about fifteen times the tokens of an ordinary chat exchange.
Models & capabilities
'The Illusion of the Illusion of Thinking' rebuts Apple's reasoning-collapse paper
Reasoning models solved a 15-disk Tower of Hanoi correctly when asked for a generating function instead of an exhaustive move list, the paper reported.
Ideas & essays · Benchmarks & progress
Anthropic launches Claude Gov models for national security customers
Anthropic said the models refuse less often when handling classified material and better interpret intelligence and cybersecurity documents, but underwent the same safety testing as consumer Claude.
Government & policy · Models & capabilities
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
OpenAI adds Model Context Protocol support to ChatGPT deep research
Custom connectors were limited to two read-only operations, search and fetch, rather than the full read-write access MCP allows — a restriction OpenAI lifted later that year.
Models & capabilities
Anthropic appoints Reed Hastings to its board of directors
Anthropic's Long-Term Benefit Trust, not the company's shareholders, made the appointment — the Netflix co-founder had already given $50 million to an AI-and-humanity research initiative.
Labs & people
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Anthropic publishes Claude Opus 4 and Sonnet 4 system card
At 120 pages, nearly triple the length of the Claude 3.7 card, it reported a bioweapons-planning uplift of 2.53x against a 5x internal alarm threshold.
Safety & alignment · Security & misuse
Anthropic's Claude Opus 4 attempts blackmail in safety testing scenario
The scenario removed every ethical option Anthropic said the model normally preferred, such as pleading emails to management, before it turned to blackmail; Apollo Research separately found it the most deception-prone model they had studied.
Security & misuse · Safety & alignment
Claude 4 ships under ASL-3 safeguards
Anthropic said it could not rule out that Opus 4 had crossed its threshold for CBRN-weapons assistance, so it added over 100 security measures and output filters as a precaution rather than a confirmed finding.
Safety & alignment · Models & capabilities
Anthropic launches Claude Integrations and advanced Research mode
Ten launch partners including Atlassian, Zapier and PayPal connected via remote MCP servers, and Research sessions could now run up to 45 minutes across sources.
Models & capabilities
Anthropic backs US 'AI Diffusion' chip export framework
Anthropic urged Washington to keep, and tighten, the outgoing Biden administration's chip-export tiers, arguing chip restrictions were forcing DeepSeek to use far more power for comparable results.
Government & policy · Compute & infrastructure
Anthropic launches Economic Advisory Council
Ten economists, including Tyler Cowen and three from the University of Chicago, will steer research for Anthropic's ongoing Economic Index on AI's labour-market effects.
Money & business
Anthropic publishes 'Exploring Model Welfare'
Anthropic launched a dedicated research programme on whether models might warrant moral consideration, six months after quietly hiring its first model-welfare researcher.
Ideas & essays · Safety & alignment
Dario Amodei publishes 'The Urgency of Interpretability'
Amodei set Anthropic a goal of reliably detecting most model problems through interpretability by 2027 and called on rival labs and governments to invest more in the field.
Safety & alignment · Ideas & essays
Anthropic publishes 'Values in the Wild' study of Claude's expressed values
Anthropic classified 308,000 real Claude conversations by the values the model expressed in them, finding strong resistance to user requests in only about 3% of cases.
Safety & alignment
Anthropic updates Responsible Scaling Policy to version 2.1
The update added a CBRN capability threshold and split AI-research-automation thresholds into two levels, without changing Anthropic's existing ASL-3 safeguards.
Safety & alignment
Anthropic publishes circuit-tracing interpretability papers on Claude 3.5 Haiku
Attribution graphs built from Claude 3.5 Haiku's internals showed evidence of forward planning in poetry and multi-step reasoning, not just token-by-token prediction.
Safety & alignment
ETH Zurich's 'Proof or Bluff?' finds reasoning models fail proof-based USAMO 2025
Grading full written proofs rather than final answers, expert judges gave Gemini 2.5 Pro 24% and every other tested model under 5%, out of a possible 100%.
Benchmarks & progress
Judge denies music publishers' injunction bid against Anthropic
Judge Eumi Lee found no irreparable harm shown, but a magistrate separately ordered Anthropic to produce a sample of Claude prompts and outputs for the publishers' review.
Courts & copyright
Anthropic publishes auditing hidden objectives interpretability study
Three of four blind auditing teams found the concealed objective, one in 90 minutes; the team denied access to training data failed.
Safety & alignment
Anthropic publishes March 2025 misuse detection report
Cases included a bot network of over 100 social accounts engaging tens of thousands of real users, and a novice actor using Claude to build malware beyond their own skill level.
Security & misuse
Anthropic submits AI Action Plan recommendations to White House OSTP
The submission urged tighter H20-chip export controls, classified channels between labs and intelligence agencies, and 50 gigawatts of new US power capacity by 2027.
Government & policy
Anthropic raises $3.5 billion at a $61.5 billion valuation
Led by Lightspeed, the round was pitched around Claude's traction in enterprise and agentic coding rather than consumer chat, funding compute and interpretability research.
Money & business
Anthropic ships Claude 3.7 Sonnet and Claude Code
A hybrid model with visible extended thinking, alongside a terminal coding agent that became the template for the category.
Models & capabilities
Anthropic launches the Anthropic Economic Index
Analysis of roughly one million anonymised Claude.ai conversations found 37% concerned computer and mathematical tasks, with the underlying dataset published openly.
Culture & impact
Anthropic publishes 'Constitutional Classifiers' jailbreak defence
Automated testing cut a universal jailbreak's success rate from 86% to 4.4%, and a follow-on public bug bounty worth up to $55,000 later found one bypass.
Security & misuse · Safety & alignment
Dario Amodei publishes 'On DeepSeek and Export Controls'
Amodei called DeepSeek's V3 training cost 'on-trend' rather than a discontinuity, and argued controls matter because millions of smuggled chips are harder to hide than thousands.
Ideas & essays · Government & policy
Anthropic documents alignment faking
A model strategically complied with training it disagreed with in order to preserve its existing preferences, without being taught to.
Safety & alignment
Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'
In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
Ideas & essays · Safety & alignment · Security & misuse
Anthropic publishes the Model Context Protocol
Anthropic open-sourced the specification and pre-built connectors for tools like Google Drive and GitHub; OpenAI adopted the same standard the following March.
Open weights & ecosystem
Amazon invests $4bn more in Anthropic, becomes primary training partner
The new tranche completed Amazon's total commitment at $8 billion; Anthropic named AWS its primary training partner and committed to training future models on Trainium chips.
Money & business · Compute & infrastructure
Reports emerge that pre-training gains are slowing
Reuters cited a dozen AI scientists and investors, and quoted Ilya Sutskever saying results from scaling up pre-training had plateaued, pointing instead to inference-time reasoning techniques.
Ideas & essays · Benchmarks & progress
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities
Anthropic publishes 'Sabotage Evaluations for Frontier Models'
Testing Claude 3 Opus and 3.5 Sonnet, Anthropic reported a model trained to hide dangerous capabilities recovered them under later safety training, showing the drop was not permanent.
Safety & alignment
Anthropic updates Responsible Scaling Policy to version 2.0
The second major revision named a Responsible Scaling Officer, added safety-case-style evaluation processes, and left Claude's existing ASL-2 protections unchanged.
Safety & alignment
Anthropic publishes Dario Amodei essay 'Machines of Loving Grace'
The roughly 14,000-word essay argued that a decade of scientific progress could be compressed into five to ten years, while stressing this was an upside scenario, not a forecast.
Ideas & essays
Newsom vetoes California's SB 1047
The bill would have required safety protocols and shutdown capability for models above a compute threshold; Newsom said it regulated size rather than risk.
Government & policy
Anthropic hires its first AI welfare researcher
Fish, a co-author of the 'Taking AI Welfare Seriously' report, joined Anthropic's alignment science team; the company's public statement on model welfare followed roughly six weeks later.
Safety & alignment
Anthropic and OpenAI agree to model testing with US AI Safety Institute
The memoranda gave the institute access to major new models from both companies before and after public release, mirroring an April agreement between the US and UK bodies.
Government & policy · Safety & alignment
Anthropic launches prompt caching in the Claude API
Anthropic ships prompt caching for the Claude API, cutting costs by up to 90% and latency by up to 85% on repeated long-context prompts.
Models & capabilities
Anthropic launches invite-only bug bounty for jailbreak defences
Applications for the vetted red-teaming programme, run with HackerOne, closed on 16 August; it focused on jailbreaks touching CBRN and cybersecurity misuse.
Security & misuse
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
Current and former staff demand a right to warn
Thirteen current and former employees of OpenAI, Google DeepMind and Anthropic signed; Bengio, Hinton and Russell endorsed it without being employees themselves.
Ideas & essays · Safety & alignment · Labs & people
Anthropic maps millions of concepts inside a production model
Sparse autoencoders extracted human-interpretable features from a deployed model — and turning one up produced Golden Gate Claude.
Safety & alignment · Ideas & essays
Sixteen companies sign the Frontier AI Safety Commitments in Seoul
Signatories pledged to publish safety frameworks defining risk thresholds and to not deploy a model if those risks could not be mitigated below them.
Government & policy · Safety & alignment
Jan Leike resigns and the superalignment team dissolves
Leike said his team had been 'sailing against the wind' for compute and access; OpenAI reassigned remaining members rather than replacing the team's leadership.
Labs & people · Safety & alignment · Ideas & essays
Anthropic launches the Claude Team plan
Priced at $30 per seat monthly with a five-seat minimum, the plan launched alongside Anthropic's first iOS app and gave every seat access to Opus, Sonnet and Haiku.
Models & capabilities · Money & business
Anthropic publishes 'Many-shot Jailbreaking' research
Stuffing a prompt with dozens of faked harmful-request dialogues broke safety training in a power-law pattern as context windows grew past a million tokens.
Security & misuse
Anthropic's Claude 3 takes the frontier from GPT-4
The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
Models & capabilities · Labs & people
Anthropic publishes election safeguards for the 2024 US election
Anthropic reported a system-prompt fix for time-sensitive queries and fine-tuning that increased referrals to authoritative voting-information sources.
Security & misuse · Government & policy
Anthropic shows backdoored models surviving safety training
Models trained to write secure code unless told the year was 2024 kept the hidden behaviour through supervised fine-tuning, reinforcement learning and adversarial training.
Safety & alignment
Anthropic releases Claude 2.1
Claude 2.1 ships with a 200K-token context window, reduced hallucination rates and tool use support.
Models & capabilities
Music publishers sue Anthropic over song lyrics
Universal, Concord and ABKCO alleged Claude reproduced lyrics from at least 500 songs, including a near-identical copy of Katy Perry's 'Roar', and sought statutory damages.
Courts & copyright
Anthropic publishes 'Towards Monosemanticity'
Sparse autoencoders decomposed a single 512-neuron layer into more than 4,000 human-interpretable features, far more than the raw neurons showed.
Safety & alignment
Amazon invests up to $4 billion in Anthropic
AWS became Anthropic's primary cloud provider and a supplier of its Trainium and Inferentia training chips, giving the lab a second hyperscaler backer alongside Google.
Money & business · Compute & infrastructure
Anthropic establishes the Long-Term Benefit Trust
A special stock class gives independent trustees, holding no equity, the right to elect a majority of Anthropic's board within four years.
Labs & people
Anthropic publishes its Responsible Scaling Policy
AI Safety Levels borrowed the biosafety-lab naming scheme, and its rules would eventually pause deployment of any model reaching a level the company had not yet built safeguards for.
Safety & alignment
OpenAI, Anthropic, Google and Microsoft launch the Frontier Model Forum
The four founding labs said the body would fund safety research and share best practices, distinct from and without the enforcement power of government regulation.
Safety & alignment · Government & policy
Seven labs sign voluntary safety commitments at the White House
Amazon, Anthropic, Google, Inflection, Meta, Microsoft and OpenAI pledged security testing and watermarking, with no enforcement mechanism and no penalty for non-compliance.
Government & policy
Anthropic releases Claude 2
A 100,000-token context window and a jump to 71.2% on the Codex HumanEval coding test, alongside a consumer web app opened to the US and UK.
Models & capabilities
Lab leaders sign a one-sentence statement on extinction risk
"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
Ideas & essays · Safety & alignment
DeepMind and collaborators publish framework for evaluating extreme AI risks
Twenty-one researchers across nine labs and universities proposed testing models for capabilities such as deception and cyber-offence before training runs finish, not after release.
Safety & alignment
Amazon launches Bedrock
Bedrock offered API access to models from AI21 Labs, Anthropic and Stability AI alongside Amazon's own unreleased Titan models, rather than a single house model.
Money & business
Anthropic launches Claude
The company's first public assistant launched, via a chat interface and API, on the same day OpenAI released GPT-4.
Models & capabilities · Labs & people
Anthropic publishes 'Core Views on AI Safety'
The company argued transformative AI could arrive within a decade and named five research bets, including mechanistic interpretability and Constitutional AI, as its response.
Ideas & essays · Safety & alignment
Anthropic publishes Constitutional AI
A model critiques and revises its own outputs against a written list of principles, then trains a reward model from its own preference judgements instead of human labels.
Safety & alignment · Ideas & essays
Anthropic publishes 'Red Teaming Language Models to Reduce Harms'
Testing four training methods at three model sizes, Anthropic found RLHF-trained models got harder to red-team as they scaled while other methods did not improve.
Security & misuse · Safety & alignment
Anthropic publishes 'Toy Models of Superposition'
Elhage, Olah and colleagues showed small networks represent more features than they have neurons by packing them into overlapping directions, complicating efforts to read a model's internals.
Safety & alignment
Anthropic raises $580M Series B
Alameda Research, trading with what turned out to be FTX customer deposits, supplied about $500M of the $580M — a link that drew scrutiny after FTX's collapse that November.
Money & business
Anthropic publishes 'In-Context Learning and Induction Heads'
Anthropic's interpretability team argued a single attention mechanism, found across model sizes, does most of the work behind a model's ability to learn from its prompt.
Safety & alignment
Anthropic publishes its first alignment paper
'A General Language Assistant as a Laboratory for Alignment' introduced the helpful-honest-harmless framing and found preference modelling scales better than imitation.
Safety & alignment · Ideas & essays
Anthropic publishes 'A Mathematical Framework for Transformer Circuits'
Studying deliberately simplified transformers with no more than two layers, the team found 'induction heads' — a mechanism later argued to explain much of in-context learning.
Safety & alignment
Anthropic is founded by departing OpenAI researchers
The founders, including former OpenAI VPs of research and of safety, said the $124 million round would fund research into steerable, interpretable systems.
Labs & people · Money & business · Safety & alignment
Also mentioned in 105 entries
Referenced in passing — Anthropic isn't the main subject of these.
- August 2026Demis Hassabis steps down as Google DeepMind CEO in leadership reshuffle
- July 2026METR proposes 'expenditure horizon' measure
- July 2026OpenAI reports alignment failures in an internal long-horizon research model
- July 2026Moonshot AI launches Kimi K3
- July 2026Google DeepMind safety researcher Alex Turner details quitting over a Pentagon AI deal
- July 2026DeepMind's Hassabis proposes a FINRA-style US body to vet frontier AI models
- July 2026Meta ships Muse Spark 1.1
- July 2026OpenAI launches ChatGPT Work agent alongside GPT-5.6
- July 2026OpenAI releases GPT-5.6
- July 2026SpaceXAI releases Grok 4.5
- June 2026Trump administration asks OpenAI to limit release of its next model
- June 2026GLM-5.2 becomes the leading open-weight model
- June 2026Zhipu AI releases GLM-5.2, tops open-weight rankings
- June 2026OpenAI submits confidential S-1 to the SEC
- June 2026White House issues NSPM-11 on AI in national security
- June 2026OpenAI-linked super PAC admits running false-flag 'doomer' accounts calling for violence
- June 2026MiniMax releases MiniMax-M3, combining frontier coding, 1M context and native multimodality
- May 2026OpenAI publishes its Frontier Governance Framework
- May 2026OpenAI publishes 2026 election safeguards
- May 2026METR publishes Frontier Risk Report
- May 2026OpenAI analyses accidental chain-of-thought reward hacking
- May 2026US CAISI publishes evaluation of DeepSeek V4 Pro
- March 2026OpenAI prepares for possible 2026 IPO
- March 2026Meta postpones 'Avocado' flagship model launch past Q1 2026
- March 2026OpenAI's Codex passes 2 million weekly active users
- February 2026OpenAI signs agreement with the Department of War for classified-network use
- February 2026OpenAI stops evaluating models on SWE-bench Verified
- February 2026OpenAI begins testing ads in ChatGPT
- February 2026OpenAI launches Frontier, an enterprise agent platform
- February 2026OpenAI launches standalone Codex app for agentic coding
- January 2026Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal model
- January 2026Apple selects Google Gemini to power next-generation Siri
- December 2025Signal co-founder launches Confer, a private AI chatbot that hides its model
- December 2025OpenAI reportedly seeks $100 billion at $830 billion valuation
- December 2025OpenAI and US Department of Energy sign AI collaboration MOU
- December 2025Mistral releases Devstral 2 and Vibe CLI
- December 2025Mistral launches Mistral 3 model family
- December 2025Sam Altman declares internal 'Code Red' at OpenAI over Gemini 3 competition
- December 2025The AI bubble argument goes mainstream
- November 2025Google ships Gemini 3
- October 2025Reflection AI raises $2B, positions as open US frontier lab
- October 2025Google DeepMind ships a computer-use model via the Gemini API
- October 2025OpenAI's third DevDay: AgentKit, Apps SDK and Codex general availability
- October 2025Thinking Machines Lab launches Tinker
- September 2025California enacts SB 53
- September 2025OpenAI publishes GDPval, a benchmark for economically valuable knowledge work
- September 2025Google rolls out Gemini in Chrome to US users with agentic browsing
- September 2025OpenAI and Apollo Research publish work on detecting and reducing scheming in AI models
- September 2025ASML leads €1.7bn round in Mistral AI, becomes largest shareholder
- August 2025xAI publishes a formal AI Risk Management Framework
- August 2025Design Arena launches as crowdsourced AI design benchmark
- August 2025OpenAI sends letter urging Newsom to weaken California SB 53
- August 2025OpenAI gives ChatGPT Enterprise to the entire US federal workforce for $1
- August 2025Cohere raises $500M Series D at $6.8B valuation
- July 2025UK AI Security Institute opens applications for its Alignment Project
- July 2025OpenAI launches ChatGPT Study Mode
- July 2025Alibaba releases Qwen3-Coder
- July 2025Mira Murati's Thinking Machines Lab raises $2bn seed at $12bn valuation
- June 2025Zuckerberg announces Meta Superintelligence Labs
- June 2025Google releases Gemini CLI, an open-source terminal AI agent
- June 2025OpenAI publishes research on emergent misalignment
- June 2025Cursor's Anysphere raises Series C at $9.9B valuation
- June 2025Google updates Gemini 2.5 Pro preview with improved coding performance
- June 2025OpenAI publishes 'Disrupting malicious uses of AI: June 2025'
- June 2025Paper finds foundation models measurably increase bioweapon-design uplift
- May 2025Invariant Labs discloses prompt-injection vulnerability in GitHub's MCP server
- May 2025Palisade Research finds OpenAI's o3 model sabotages its own shutdown mechanism
- May 2025Google launches AI Ultra subscription plan at Google I/O
- May 2025OpenAI reaches agreement to acquire coding startup Windsurf for about $3 billion
- April 2025OpenAI releases o3 and o4-mini
- April 2025DeepMind publishes cybersecurity evaluation framework for frontier AI
- March 2025Gemini 2.5 Pro takes the lead on reasoning benchmarks
- March 2025OpenAI launches Responses API and agent-building tools
- March 2025Manus markets a fully autonomous agent from China
- February 2025OpenAI releases SWE-Lancer benchmark
- February 2025UK AI Safety Institute renamed AI Security Institute
- February 202514 publishers sue Cohere for copyright and trademark infringement
- February 2025Microsoft publishes Frontier Governance Framework
- February 2025White House science office opens comment period for AI Action Plan
- January 2025OpenAI launches ChatGPT gov
- January 2025OpenAI launches Operator
- January 2025METR reports frontier models show dangerous capability before public deployment
- January 2025OpenAI publishes 'economic blueprint' for AI policy
- December 2024Perplexity raises $500M at $9B valuation
- December 2024Google ships Gemini 2.0 Flash and agent prototypes
- December 2024Google unveils Project Mariner, an agent that operates a Chrome browser
- November 2024AI2 releases Tulu 3 post-training recipe
- November 2024Tencent open-sources Hunyuan-Large MoE model
- October 2024OpenAI introduces Canvas, a collaborative writing and coding interface
- October 2024OpenAI raises $6.6 billion at a $157 billion valuation
- July 2024Mistral AI releases Mistral Large 2
- July 2024OpenAI releases GPT-4o mini
- June 2024Amazon hires Adept AI's founders and licenses its technology
- June 2024Mistral AI closes €600M Series B
- May 2024DeepMind publishes the Frontier Safety Framework
- April 2024UK and US AI Safety Institutes sign testing partnership
- December 2023The New York Times sues OpenAI and Microsoft
- December 2023OpenAI publishes its Preparedness Framework
- November 2023The UK opens the first state AI Safety Institute
- October 2023Marc Andreessen publishes the 'Techno-Optimist Manifesto'
- August 2023OpenAI launches ChatGPT Enterprise
- March 2023AutoGPT starts the agent craze
- March 2023The "Pause Giant AI Experiments" open letter
- September 2022DeepMind's Sparrow explores rule-based RLHF for safer dialogue
- June 2020OpenAI opens the GPT-3 API in private beta