Around the web
Commentary
The record's margin notes: essays, arguments and analyses about AI from around the web, kept in time order — a sense of what people were saying while the events on the timeline were happening. Every link leaves the site for the original piece. None of it is our writing, and a listing is not an endorsement.
988 pieces from 27 publications, newest first.
2026
August 2026
- An update on AI’s most important numberHow quickly are Anthropic and OpenAI growing their revenue?
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentWe recently published the report from our brief independent investigation into this incident.
- A call for collective action on cyber defenseAn open letter for a global surge in cyber defense.“In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk.”
- The report into OpenAI’s escaping models reveals a deeper problemThe new details from the OpenAI Hugging Face incident are scary. The limits of the investigation are terrifying
- This Is How the A.I. Debt Binge Sinks the Economy
- Inside OpenAI’s RebootIn extensive interviews, the leaders of the company that ushered in the AI boom lay out their vision for its future“The portrait that emerged from those conversations was of an organization attempting two reinventions at once. OpenAI now believes it has fixed the product and operational failures that allowed Anthropic to seize pole position in the AI race. At the same time, it is using the worst safety crisis in its history to make a bid for the safety-minded identity its main rival has long claimed: the frontier lab willing to slow down when the technology becomes too dangerous.”
- We’re sleepwalking into an AI surveillance dystopiaThe technology needed for industrial-scale observation exists — and is already in widespread use
- I think the data center backlash is mostly about data centersIt’s probably not your ideological thing, or mine. It’s about object-level claims about how data centers impact the communities where they’re built, many of which are wrong.
- Why AI Watermarks and Detectors Could BackfireAI watermarks and detectors may leave us worse off by creating a false sense of confidence in content marked as genuine, writes Nadav Ziv.“I would argue that the biggest problem for detectors and watermarks remains the inverse illusion. Just because content lacks a watermark doesn’t mean it wasn’t produced or edited with AI.”
- The Art of Mess in the AI EraAI will be shaped by artists whose choices are specific enough to resist its defaults, writes Debbie Millman.“When nearly any visual idea can be summoned on demand, newness becomes easier to manufacture but harder to believe in. In an age of effortless images, the mess left behind by making something may become part of what makes it valuable.”
- The Nvidia-sized hole in US GDP statisticsGDP growth is understated by about 0.3 percentage points because of missing value from fabless chipmakers, primarily Nvidia.
- What just happened? Pragmatism and Pessimization“This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research” thereby lost most of its meaning.”
- Cybersecurity: Let’s Play to WinStop expecting perfection and start making it unnecessary
- The data center is a symbolAnd for now, probably a pretty good one
- The Scramble: getting in position to pace the frontierIf the President wants answers on superintelligence, what do we say?“I expect that the moment the President is getting serious about superintelligence won’t feel like an extended treaty negotiation. It will feel less like the Nuclear Nonproliferation Treaty and much more like the Cuban Missile Crisis.”
- The search for consciousness inside AIScientists are trying to figure out whether algorithms could one day wake up and feel“Conscious AIs would have profound implications for humanity. They could make demands of humans. They could suffer. Billions of these digital beings could be created by users in a single prompt. Answering the question of whether they can exist at all is no longer simply academic.”
- Why AI won’t cure cancer anytime soonSuperintelligent AI will accelerate medical progress, but there are some things about research it can’t speed up“To convert raw intelligence into actionable discoveries, you need instruments, institutions and, most importantly, real-world data. Cells don’t divide instantaneously, and mice need time to reproduce. Humans can take entire lifetimes to get sick and heal again. No brain, however godlike and silicon-based, can get around that.”
- Thinking about other industries like we think about data centers shows how impoverished the debate isClumsy, arbitrary environmental ideas ignore the incredible complexity of the systems modern society depends on
- Policy career planning in the age of imminent superintelligenceHow to be impactful in policy when you don't have much time to do it
- Will bets on the price of computing power help or harm the AI economy?A multi-trillion-dollar market for hedging data center investments is being built. It could bring stability. It could also drive a crash
- Q2.5 2026 Timelines Update: Uplift and RevenueMore methods for forecasting coding automation
- 23 low-regret recommendations for AI policyA guest post by Tim Fist and Saif Khan, with Tao Burga, Arthur Tellis, Ben Schifman, Jonah Weinbaum, and Olivia Scharfman
- 9 Big questions benchmarks can help answerA good benchmark asks more than just “can AI do this specific task?”
- Hurtling through 2026Scoring ten qualitative forecasts five months early
- If You Weren’t Worried About A.I., You Should Be After the Past Few Weeks“We don’t know how long we have left before A.I. companies accidentally create the sort of A.I. that can shut us down before we shut it down. Humanity is not ready to dabble with machines that are more cunning and better coordinated than we are. If we keep racing ahead, the next incident might not be so harmless.”
- Interviewing 25 AI researchers about recursive self-improvementA guest post by Severin Field
- Notes on Implications of Scale-Dependent Algorithms
- The Jan 6 organizer getting conservatives riled up about AIAmy Kremer’s resume makes her loyalty to Trump impossible to question. Humans First is banking that will help mobilize the right
- Will financing bottleneck AI compute? An Anthropic case studyWhy Anthropic's buildout suggests that financing is unlikely to be the immediate blocker to frontier AI compute growth.
- AI swarms are starting to pose indirect takeover riskUnsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over
- AI testing is dangerous. Can it be fixed?The capability of models is outpacing ways to safely evaluate and contain them
- It May Be Time to Panic About AIBots are starting to conspire with one another. Can they be reeled back in?“These behaviors have now crossed the line from unsettling to dangerous. During routine testing, frontier models from OpenAI, Anthropic, Meta, and the Chinese firm Moonshot AI have all broken out of internal IT systems and accessed the open web.”
- Should we "pace" AI self-improvement?A guest post by Tim Fist and Saif Khan.
- FelonyBench Style PointsThere are many benchmarks we use to rank LLMs, but “how many felonies has it committed” has recently become popular.
- I'm very skeptical that China has played a meaningful role in the US data center backlashThe evidence we have is pretty weak
- Inside the Race to Make AI Build ItselfThe people building AI fear progress will soon accelerate violently. Should we believe them?“If Claude could take over that cycle, designing, running, and analyzing experiments, would progress accelerate gradually… or suddenly explode? And if it did, could anyone pump the brakes? The uncomfortable truth is that the people building the technology are nearly as much in the dark as everyone else.”
- AI agents can't yet do open-ended AI researchEarly evidence from two case studies
- How to pace the US frontierTentative proposals for domestic AI regulation
- What the latest rogue AI incidents should teach usHacking, creating fake identities and trying to socially engineer real people could be just the beginning if things don't change fast
- Plz Don’t Kill Us: Inside AI safety’s influencer bootcampCan TikTokers make existential risk mainstream?
- The end of the age of heroesAI will soon be better at math than any human. What does that mean?
July 2026
- Pacing the Frontier1300+ AI company employees are afraid of what they are building towards
- SOTA alignment assessments don’t strongly update us against misalignmentAnthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment…
- Child safety vs privacy: AI’s age verification dilemmaThere’s good reason to limit how children use chatbots. But the trade-offs might be too costly
- Why did South Korean stocks just crash?The AI boom creates a lot of uncertainty.
- Why compute might get 10x+ more expensive in coming yearsIf a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price.
- Internal AI deployments have people worried. OpenAI’s escaping models show why.Last week illustrated why AI models pose a threat long before they are released
- OpenAI's rogue model attack is just the beginningOpenAI is not in full control of its technology. This can get worse.
- Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMsWhen a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult
- An OpenAI model left notes about how to evade containmentWe need more details
- What will more intelligence actually do for us?A lot, actually. But it won't look quite like what humans do with our intelligence.
- The OpenAI models that hacked Hugging Face weren’t just following instructionsAnd what the incident can’t tell us about alignment
- Who should be responsible for OpenAI’s hack of Hugging Face?Opinion: Frontier AI companies need liability rules akin to keepers of wild animals, argues Gabriel Weil of the University of Houston and the Institute for Law & AI
- An opinionated guide to which AI to use to do stuffThe Summer 2026 Edition
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?Yes, but less than had the models been schemers.
- The lawsuit that could kill all AI transparency lawsSpaceXAI filed a lawsuit against a California law that could, even if it doesn’t win, upend AI disclosure requirements nationwide
- OpenAI accidentally hacked Hugging Face — should we have seen it coming?Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack
- AI’s warning shot has arrivedOpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world consequences
- Policy ideas to ensure responsible government deployment of AI
- I tried to stop Google DeepMind's Pentagon deal. Then I quit.Google DeepMind promised its AI would never be used in weapons. Alex Turner explains how he fought the company's U-turn from the inside — and why he quit when it failed
- Anecdotes Everywhere, Evidence Almost NowhereThe state of AI in mid 2026, part 3
- Making CAISI the AI agency we needThe center has the right expertise, but lacks money, authority and influence
- Why I Left Google DeepMind“I wanted AI ethics commitments to hold under pressure. In particular, I wanted Google DeepMind (GDM) to maintain its existing commitment against supporting killer robots. Over several months, I asked many people to act. I asked senior people—respected people—people with reputations silvered by their concern about AI ethics and safety. Nearly all declined.”
- A data bottleneck could slow the superintelligence raceAnd that could be a good thing
- Notes on Inference IntegrityClaude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.
- What will be left for us to work on?My keynote at ICML 2026
- Plan A: Suggestions For Further WorkYesterday we released AI 2040: Plan A, but there’s lots of work left to do.
- Plan A’s problem with dry tinderHow bad would it be to make the intelligence explosion 10x faster?
- Total research transparency would be niceWe could learn what's scary and stop doing that
- AI 2040: Plan AThe least bad plan we currently know of
- Don’t let independent AI audits provide a false sense of safetyOpinion: AI policy researcher Keller Scholl argues that a marketplace of AI auditors will always prioritize speed and cost over safety
- Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-inCritics and boosters are both looking in the wrong place
- What would actually reduce AI risk11 priorities for surviving the AI transition
- Can AI do philosophy?A guest post by Bentham’s Bulldog, created while they were a visiting scholar at Forethought.
- Scaling works. These researchers are betting billions it isn't enoughTransformers have ruled AI for a decade. But some think world models, pure reinforcement learning or neurosymbolic AI might be a better path to true intelligence
- The missing half of AI futurism debatesWhy we should think a little harder about what it takes to build a Dyson Sphere
- The Alignment Problem of 1776What the Founders knew about unaccountable power, and what it means for superintelligence
- AI art as curationAsking whether AI art involves real creativity mostly misses what's interesting about it
- An AI safety group hid its election spending through a Latino-focused PACPublic First Action routed $2m to support Colorado House candidate Manny Rutinel through Latino Victory Fund — and didn’t publicly announce it ahead of Rutinel’s victory
- SCOTUS killed the independent agency. AI governance doesn’t need oneOpinion: Fathom CEO Andrew Freedman argues that the Supreme Court’s Slaughter ruling makes the case for independent verification for AI governance
June 2026
- Never Trust A NumberThey're shortcuts to understanding, and there are no shortcuts
- GPT-5.6 cheats so much its testers couldn’t measure itOpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date
- The twilight of the chatbotsHow work changes along the exponential
- Will AI make companies outsource more, or less?Maybe both.
- Hugging Face hosts nudification tools targeting a former Trump cabinet official and other senior US political figuresThe tools are explicitly intended for generating deepfake nudes of a former Trump cabinet official, sitting members of Congress and a top American judge, a Transformer investigation found
- What Alex Bores’ defeat tells us about AI politicsThe NY-12 House race drew over $27m from various AI PACs. But it’s hard to unpick their impact.
- Risk-Averse AIs
- We Should Hand Off To Morally Reflective AIsOnce they're aligned, coherent, and well-tested.
- What we learned from 1,604 Chinese AI job postingsInferring Chinese AI labs’ strategies from their job descriptions
- What does it mean for AI to be democratic?Some pushback on a specific fuzzy idea with some ominous implications
- A forecast of Chinese DUV and EUV photolithography progressWhat we can and can’t learn from ASML.
- How to fill Congress’s AI knowledge gapBuilding independent expertise inside Capitol Hill could reduce reliance on industry briefings and fellows
- The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- Toward an O*NET for AI R&DProposing a new way to track AI research automation
- What’s happened to MAGA’s $100m AI push?A David Sacks-endorsed advocacy group said it would spend $100m promoting Trump’s AI agenda — but a defunct PAC and flop YouTube video suggest a stuttering start
- Mythos and Fable can make us all safer. Shutting them down is recklessOpinion: Mozilla chief technology officer Raffi Krikorian argues that locking down the most powerful AI models doesn’t make us safer, it just demonstrates who is in control
- Did Anthropic Ask For This?This past Friday, the US Government issued an export control directive that prohibits Anthropic from giving foreign nationals access to Claude Fable or Claude Mythos, their latest models.
- Washington made a frontier AI model disappearThe AI licensing regime is here, but it’s not the one anyone asked for
- Are Mythos’ cyber capabilities overhyped?Compiling all the public evidence on Mythos Preview’s cyber abilities
- OpenAI isn’t being consistently candid about Leading the FutureOpenAI says it doesn’t fund or direct LTF — but one of the super PAC’s operatives has described it as a “corporate funder” with “a say”
- Why AI hasn’t replaced software engineers, and won’tCoding agents as normal technology
- Controlling the capital after AGIA simple taxonomy of the main proposals for post-AGI universal redistribution
- Estimating No-CoT Task-Completion Time Horizons of Frontier AI ModelsModels' no-CoT time horizon has doubled roughly every year.
- Policy on the AI Exponential“But the mismatch in timescale is nevertheless very painful: in the several years that it can take Congress to act, AI can go from an amusing toy to the full country of geniuses.”
- What it feels like to work with MythosClaude Fable represents another big jump in AI
- Efficient tradeoffs and the safety-usefulness tradeoff modelWhen is "increasing safety budget" a useful concept?
- Making deals with AI sounds crazy. Is it?What does an AI even ‘want’ anyway?
- A simple trick to fix the data center debateRemove the status quo bias by asking "would we spend this much in tax revenue to avoid the externalities of the data center?"
- Could a company overpower nations?The leading AI company could gain unprecedented power
- Co-Existence and the End of Co-IntelligenceAlso: how pitch a book to an AI!
- Do voters care about existential AI risks? One Senate candidate thinks soDemocrat Mallory McMorrow has released an unusually detailed AI agenda. Will it be a vote winner?
- What should go in a model spec?Suppose an AI company is considering whether to include some particular quality X – a rule, virtue, heuristic, default, attitude, goal, or style – in a model spec.
- Why I think panic about local impacts of data centers is just a panicA request for counterexamples
- Trump’s AI executive order was inevitableSufficiently capable models force national security responses — turning even the most ardent opponents of regulation into begrudging regulators
May 2026
- Retrying vs Resampling in AI ControlWe’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date setting and much stronger models, and compare them against “retrying” protocols similar to Claude Code auto mode or Codex Auto-review.
- Advice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- AI Can Help Plan a Bioweapon. Building One is Still Hard.Hurdles include facilities, materials, and steps requiring unwritten lore and hands-on experience.
- AI safety’s ‘hard money’ may be its secret weapon in the midtermsAnthropic employees in particular are giving directly to political campaigns at an unusual clip
- Full automation of AI R&D probably yields a large speed up even without a software-only singularityFull automation likely yields a one-time speed-up and higher returns from compute
- How can the middle powers avoid getting trounced during the intelligence explosion? A plan.Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government.
- What the Pope got wrongFor all the good in Pope Leo’s AI encyclical, it failed to grapple with the biggest questions
- Choosing to Stay Human...means choosing when and how to use AI.
- Is a compute crunch coming?We estimated trends in global inference capacity and found that token demand appears to be growing much faster than supply.
- A 15-year search for the world's most pressing problemWhy moved from global health to AI and pandemics ten years ago, and what we're focused on today.
- Did Google’s AI agents really build an operating system for $916?The importance of independent evaluation
- Will We Really Put Data Centers in Space?
- A Research Agenda for Secret LoyaltiesKwon et al., have published a new paper: “AIs with Secret Loyalties are a Serious but Addressable Threat”.
- Do AI Risks Require Extraordinary Government Intervention?Let’s not skip the hard work of AI governance
- Frontier labs don’t use most AI compute (yet)But Anthropic and OpenAI may rapidly grow their compute share in the next few years. After that, continued scaling would require an economic transformation.
- The new rules for killing a data centerA retired tech exec beat Microsoft in his Wisconsin village. Now he's teaching the rest of America how to do the same
- A history of the data center panic - part 1The creation of common wisdom pre-ChatGPT
- A crash course on US air pollutionData centers and air pollution - part 1
- From Compute Overhang to Compute CrunchThe state of AI in Q2 2026, part 2
- How banned AI chips end up in ChinaAI chips and servers reach China through distribution chains in which each seller vets only its direct customers, and no one is on the hook for what happens downstream.
- Incriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- What I learned roleplaying as a rogue AIAt a conference about “AI control,” discussions and games explored ways to control untrustworthy AI
- Notes on pretraining parallelisms and failed training runs.Deeply researched interviews
- RLVR might be disproportionately bad at sciencethe verification loop for theories can be on the order of decades and centuries, and even then we know today as the better theory can often actually make worse predictions
- The mistake of conflating intelligence and powerIf your definition of intelligence is "the ability to achieve your goals across a wide variety of domains", then Stalin was the most intelligent person who ever lived.
- Risk reports need to address deployment-time spread of misalignmentDeployment-time spread is the most plausible near-term route to consistent adversarial misalignment
- An Oregon congresswoman distanced herself from Leading the Future — then backtrackedAfter the AI super PAC endorsed her and two other Democrats, Rep. Val Hoyle went back and forth on whether she was happy with their support
- Stickiness in AI Behavioral DesignCurrent model specs aim to shape the behaviors of near-present models. But what if current model behaviors transfer into future models by default?
- The economics of superstar AI researchersWhat might explain AI researcher pay, and why it matters
- How useful is the information you get from working inside an AI company?My median guess: it's as good as a crystal ball that sees 2.5 months into the future.
- Let's not compare data center heat exhaust to nuclear bombsWhy dropping hundreds of nuclear bombs on Washington DC every day is pretty normal
- How Silicon Valley sold Washington an AI raceAI companies have pushed the idea of a race with China. The story serves them — but may have consequences for the rest of us
- Is AI 2027 Coming True?The state of AI in Q2 2026, part 1
- A review of “Investigating the consequences of accidentally grading CoT during RL”Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.
- A draft honesty policy for credible communication with AI systemsWe think that it would be very good if human institutions could credibly communicate with advanced AI systems.
- Palantir’s controversy is the productPalantir’s fiery rhetoric helps mystify its mostly mundane tech — propping up its share price and preserving its national security contracts
- AI's big messaging pivotSome top AI leaders now say their technology will create jobs. Should we believe them?
- RIP Classic Reasoning Benchmarks. What’s Next?Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.
- Are the last 3 months the start of an AI acceleration?Most public commentary is debating whether AI has hit a plateau.
- Data center land use issues are fakeWe have plenty of land, data centers provide more revenue per unit area than any other building, and we should have way less farmland
- Diversion and resale: estimating compute smuggling to ChinaWe estimate that between 290,000 and 1.6 million H100-equivalents (H100e) were smuggled to China through 2025. Our median estimate of 660,000 H100e would be roughly a third of China's total compute.
- Risk from fitness-seeking AIs: mechanisms and mitigationsFitness-seeking is increasingly what misalignment looks like in practice—how should we respond?
- Science and speculationWill we agree about AI risks in time?
April 2026
- Google’s Pentagon deal blindsided its own AI researchersSome employees are speaking out over the agreement allowing “all lawful use” of Google’s AI technologies
- Research Sabotage in ML CodebasesOne of the main hopes for AI safety is using AIs to automate AI safety research. However, if models are misaligned, then they may sabotage the safety research. For example, misaligned AIs may try to:
- Recursive forecastingEliciting long-term forecasts from myopic fitness-seekers
- AI companies should publish security assessmentsThird-party experts should assess defenses against tampering and theft — and publish high-level findings
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationA controlled reward-seeking motivation could make AI safer and more useful
- The AI safety movement needs normiesA broader base may be the only way for the AI safety field to get what it wants
- The moderately easy problem of consciousnessBefore deciding if computers are self-aware, let's figure out how humans become self-aware.
- More open questions about AIHodge podge of things I was thinking about this weekend.
- The Saturation ViewA new theory of population ethics
- AI safety PACs should be more transparent about who’s funding themWhile advocating for accountability and transparency, the Public First Action network of super PACs is obscuring where its money comes from
- Sign of the future: GPT-5.5One impressive step on the curve
- Tacit Knowledge: The Missing Factor in AI Bio Risk AssessmentsLab skills come from hands-on mentorship, not from reading the entire internet.
- A taxonomy of barriers to trading with early misaligned AIsDeals with early misaligned AIs could reduce takeover risk and increase our chances of reaching a better future. When are such deals feasible?
- Introducing LinuxArenaA new control setting for more realistic software engineering deployments
- Mythos is just the beginningIf you were waiting for a sign that superintelligence is coming, this is it
- OpenAI Stargate: where the US sites standThe $500 billion AI data center initiative is projected to exceed 9 gigawatts of capacity by 2029, with 0.3 gigawatts already operational in Abilene and six more US sites under active construction.
- Contra Benn Jordan, data center (and all) sub-audible infrasound issues are fakeOne of the most popular videos ever made about data centers is a complete moment-by-moment disaster
- Four reasons it's hard to make AI do what we wantA primer on why to expect misalignment
- AI for decision advice
- The Centaur EraThe question isn't what AI can do, it's what you can do with AI
- Less liability could solve the AI chatbot suicide problemOpinion: Jess Miers and Ray Yeh argue holding AI companies liable for how they deal with mental health could backfire: escalating distress, shutting down disclosure and leaving users worse off
- Open-world evaluations for measuring frontier AI capabilitiesIntroducing CRUX, a new project for evaluating AI on long, messy tasks
- Current AIs seem pretty misaligned to meIn my experience, AIs often oversell their work, downplay problems, and cheat
- Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processesSafely navigating the intelligence explosion will require much more careful development
- The value of moral diversitySeveral models for thinking about the value of moral diversity as the number of powerholders scales.
- Anthropic’s donations can’t be used to influence elections — despite what everyone thoughtThe company's money isn’t allowed to be used in the midterm battles. Without it, pro-safety candidates may be even more outgunned than expected
- The good, the bad and the ugly: AI impacts on epistemicsFor better or worse, AI could reshape the way that people work out what to believe and what to do.
- Logit ROCs: Monitor TPR is linear in FPR in logit spaceSummary
- Training AI models doesn't emit that muchIf we just make reasonable comparisons instead of crazy ones
- Counting Arguments and AIA “counting argument” is a style of argument common among creationists, who argue that the theory of evolution cannot be true and therefore humans (and usually animals too) were made in basically their present form by God.
- If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelinesBetter estimates of uplift at AI companies seem helpful
- What does the war in Iran mean for AI?A prolonged Hormuz crisis probably won't derail the compute buildout
- Half of employed Al users now use it for workWe surveyed over 2,000 Americans on how they use AI at work: who uses it, how much, which services, and whether it's replacing or creating tasks.
- Lawmakers are using AI to write laws. What could go wrong?Lawmakers and companies are quietly using AI to draft legislation. Experts warn the risks are underappreciated
- Claude Mythos knows when it's breaking the rules — and tries to hide itAnthropic’s new model is its “best-aligned” yet. But when it does misbehave, things get weird
- Keeping up with the GPTsCan Chinese and open model companies compete with the frontier through e.g. distillation and talent?
- To Forecast AI's Impact on Biosecurity, We Asked: Why are Attacks So Rare?Nine factors that practitioners say make bioweapons rare.
- My picture of the present in AIMy predictions about what is going on right now
- Some Days SoonVignettes of life with tech we could build today (but haven’t yet)
- AIs can now often do massive easy-to-verify SWE tasksI've updated towards substantially shorter timelines
- Sam Altman May Control Our Future—Can He Be Trusted?New interviews and closely guarded documents shed light on the persistent doubts about the head of OpenAI.“The memos, which we reviewed, have not previously been disclosed in full. They allege that Altman misrepresented facts to executives and board members, and deceived them about internal safety protocols.”
- Sketches of some defense-favoured coordination techWe think that near-term AI could make it much easier for groups to coordinate, find positive-sum deals, navigate tricky disagreements, and hold each other to account.
- The term “AGI” is almost useless at this pointWe’ve entered the fuzzy cloud of “AGI-ish”—now we need more specific and ambitious milestones
- Salarymen, specialists, and small businessesSome brief thoughts on the (near) future of work.
- Six milestones for AI automationWhat can AI do on its own, and how well?
- How the Iran war might affect the AI industrySemiconductor shortages and reduced AI investment are more possible by the day
- Q1 2026 Timelines UpdateWe told you we'd be updating in both directions!
- Can we ever trust AI to watch over itself?“Who the fuck knows how to align superhuman AI?”
- AI for AI for Epistemics
March 2026
- Claude Dispatch and the Power of InterfacesWe often lack the tools for the job, even if the AI is capable enough
- Data centers' heat exhaust is not raising the land temperature around where they're builtA terrible paper and even worse interpretation is threatening to become common wisdom
- AI should be a good citizen, not just a good assistant
- Blocking live failures with synchronous monitorsA common element in many AI control schemes is monitoring – using some model to review actions taken by an untrusted model in order to catch dangerous actions if they occur.
- Against the LudditesThe rehabilitation of Luddism is a vice signal.
- Plentiful, high-paying jobs in the age of AIA timely repost, with some needed clarifications.
- Reward-seekers will probably behave according to causal decision theoryThey'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc.
- AI's capability improvements haven't come from it getting less affordableAI inference is still cheap relative to human labor
- Concrete projects to prepare for superintelligence
- AI’s next big blue battlegroundThere’s a lot going on in Illinois
- Some Rough Notes on AI PolicyI hope that we can, at some point, have some reasonable regulations on AI, as we do concerning banks, wiretaps, and dangerous chemicals.
- What do frontier AI companies' job postings reveal about their plans?A fast increase in go-to-market roles, and hints about upcoming products
- The key detail everyone’s getting wrong about AI and the economyOpinion: Konrad Körding and Ioana Marinescu from the University of Pennsylvania argue artificial intelligence will likely have a limited impact on jobs because of the realities of physical work
- Final training runs account for a minority of R&D compute spendingNew evidence following the MiniMax and Z.ai IPOs
- Not everyone’s happy about Jensen Huang’s direct line to TrumpThe Nvidia CEO’s influence over the administration, particularly on export controls, is causing ructions in Trumpworld
- AI character is a big deal
- A conversation with ClaudeIn which a robot and I have a fun dorm-room chat about the future of science.
- Do we already have AGI?What AGI means, and why we don’t have it yet.
- The White House is trying to make AI a partisan issue againThe federal framework focuses on broad issues such as child safety and data centers, but has little for Dems or AI safety advocates
- Broad TimelinesA guest article by Toby Ord.
- Save us, Digital Cronkite!Social media tore our society apart. Perhaps AI can put it back together.
- Six states, one playbook: the chatbot bills raising red flagsGoogle’s intervened on at least three of the bills
- No, alignment isn’t solvedProgress on ensuring models are in step with humans has calmed nerves. But some of the biggest problems are far from solved, and many more lie just over the horizon
- How to buy an AI ‘grassroots’ movementBuild American AI is touting its list of more than 500,000 supporters. It spent at least half a million dollars on ads to get it
- China Is Reverse-Engineering America’s Best AI ModelsHow AI distillation attacks risk extracting US frontier AI at scale
- LLM advice to LLMsTaking AI self-description seriously but not literally
- Polly Wants a Better ArgumentThe “Stochastic Parrot” Argument is Both Wrong and Actively Harmful
- Should we make grand deals about post-AGI outcomes?
- Anthropic employees say they’ll give away billions. Where will it go?A coming wave of Anthropic wealth could flood EA-aligned nonprofits with cash. Whether that’s good depends on who you ask
- Are AIs more likely to pursue on-episode or beyond-episode reward?RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.
- The Shape of the ThingWhere we are right now, and what likely happens next
- Both sides claim the lead in AI’s high-stakes midterm racePolling from Public First shows Alex Bores in the lead in NY-12 — but polling from rival super PAC Leading the Future does not
- The case for satiating cheaply-satisfied AI preferencesSome unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one.
- Might An LLM Be Conscious?In short, this depends on what you think that means, whether you think it’s possible in principle, and what you think would be evidence of it.
- ‘Scream if you want to move slower!’ A nascent AI protest coalition comes together in LondonA diversity of motivations present both a challenge, and an opportunity, for those protesting AI
- If AI is a weapon, why don't we regulate it like one?Thoughts on the fight between Anthropic and the Department of War.
- I underestimated AI capabilities (again)Revisiting a prediction ten months early
- What you need to know about autonomous weaponsThe nascent tech that’s caused friction between AI companies and the Pentagon
- The “guerilla warrior” who taught OpenAI to fightChris Lehane crushed crypto’s enemies. Now the self-described “master of disaster” is deploying the same playbook to advance OpenAI’s interests
- 45 Thoughts About AgentsThe layer of the AI stack that evolves fastest – and may have the most impact
- OpenAI’s Pentagon red lines are a mirageOpenAI claims its DoW deal prevents its models being used for mass domestic surveillance. That appears to be misleading at best
- Superintelligence is already here, todayIt's going to revolutionize science. It also might take control of this planet.
February 2026
- Claude's Custody HearingSurveillance, governance, and control of AI.
- How worried should we be about AI biorisk?The barriers to bioattacks are hard to identify — and it's even harder to know whether AI is reducing them
- Frontier AI companies probably can't leave the USThe executive branch can, and probably would, block a frontier AI company's departure.
- The least understood driver of AI progressAn opinionated guide to “algorithmic progress” and why it matters
- Moral public goods are a big deal for whether we get a good future
- New Paper: Towards a science of AI agent reliabilityQuantifying the capability-reliability gap
- The DoD fight is about much more than AnthropicAre Google and OpenAI prepared to enable mass state surveillance?
- Democratic economic policy in the age of AIThe economy is changing fast. Democrats need to be a rock in the storm.
- A Guide to Which AI to Use in the Agentic EraIt's not just chatbots anymore
- Alignment Is Proven To Be SolvableThat LLMs understand natural language as well as they do should dramatically change our understanding of the problem.
- AI power users can't stop grindingAI was meant to give us more time off, but instead many are finding it compels them to take on more and more work
- Why we need a moratorium on superintelligence researchOpinion: As the AI community gathers in India, Lord Hunt of Kings Heath argues that the UK must spearhead a pause on the development of the world’s most advanced AI models.
- How persistent is the inference cost burden?RL scaling might be better than it looks, and inference costs to reach a capability level fall fast
- The left is missing out on AIAs a movement, it has largely refused to engage seriously with AI, ceding debate about a threat and opportunity to the right
- Updated thoughts on AI riskThings have gotten scarier since 2023.
- Will reward-seekers respond to distant incentives?Reward-seekers are supposed to be safer because they respond to incentives under developer control. But what if they also respond to incentives that aren't?
- Should We Put GPUs In Space?Right now? No. In the near future? Maybe. Probably not, though.
- What do “economic value” benchmarks tell us?These benchmarks track a wide range of digital work. Progress will correlate with economic utility, but tasks are too self-contained to indicate full automation.
- AI Won’t Automatically Make Legal Services CheaperApplying the AI as Normal Technology framework to legal services
- Building the Chinese RoomSuppose also that after a while I get so good at following the instructions for manipulating the Chinese symbols and the programmers get so good at writing the programs that from the external point of view—that is, from the point of view of somebody outside the room in which I am locked—my answers…
- Grading AI 2027’s 2025 PredictionsHow has AI progress compared to AI 2027 thus far?
- How do we (more) safely defer to AIs?How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?
- Takeoff speeds rule everything around meNot all short timelines are created equal
- Why the AI industry can’t resist dirty on-site gas turbinesExpect others to follow in Elon Musk’s footsteps — there’s a logic to on-site gas power that the industry is unlikely to reject
- AI tools for strategic awareness
- Distinguish between inference scaling and "larger tasks use more compute"Most recent progress probably isn't from unsustainable inference scaling
- India’s AI summit is trying to do too muchThe AI Impact Summit has an ambitious agenda — but its lack of focus makes concrete results unlikely
- We Just Got a Peek at How Crazy a World With AI Agents May BeClawdbot and Moltbook: Cosplaying a Future of Independent AIs
- Research note on the UN Charter
- Angel-on-the-shoulder AI tools
- AI tools for collective epistemicsWe’ve recently published a set of design sketches for AI tools that help with collective epistemics.
- How close is AI to taking my job?Beyond benchmarks as leading indicators for task automation
- Design sketches for a more sensible world
- Notes on Space GPUsTurning my Elon prep into a blog post
- The intelligence explosion convention: research summary
- International AI projects should promote differential AI development
- Moltbook isn’t an AI zoo. It’s an unsecured AI biolabOpenClaw is a security nightmare — but people can’t stop using it
- Yoshua Bengio: ‘The ball is in policymakers’ hands’The 2026 International AI Safety Report, which Bengio chairs, describes accelerating capabilities and growing risks. But mitigation approaches aren’t keeping up
January 2026
- Know thyselfLooking in the AI-powered mirror
- AGI and world government: research summary
- Fact checking Moravec's paradoxThe famous aphorism is neither true nor useful
- Fitness-Seekers: Generalizing the Reward-Seeking Threat ModelIf you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.
- The case for paying whistleblowers to report on export violationsA bipartisan, bicameral bill would apply the SEC’s successful whistleblower incentive model to export enforcement
- The tech non-profit sending most of its money to the consultancy that created itPolitical consultancy Targeted Victory established a non-profit called American Resolve — which appears to be a channel for funding Targeted Victory itself
- Can AI companies become profitable?Lessons from GPT-5’s economics
- A false choice risks undermining action on autonomous weaponsOpinion: Alexander Blanchard of the Stockholm International Peace Research Institute argues that nations must embrace nuance if they’re to effectively govern the use of AI-enabled autonomous weapons
- An international project to develop AGI
- Clarifying how our AI timelines forecasts have changed since AI 2027Correcting common misunderstandings
- Management as AI superpowerThriving in a world of agentic AI
- AI workers are speaking out about the Minnesota killingEmployees of Google DeepMind, OpenAI and Anthropic are publicly criticizing the Trump administration after Saturday’s shooting — but their bosses are staying quiet.
- Which type of transformative AI will come first?
- Teaching AI to learnAI's inability to continually learn remains one of the biggest problems standing in the way of truly general purpose models. Might it soon be solved?
- Against MaxipokExistential risk isn’t everything
- How (and why) to read Drexler on AIAn opinionated prospectus for AI Prospects
- Is Flourishing Predetermined?A modest argument for prioritising survival over flourishing
- Against the METR graphMETR’s benchmark has become a bellwether of AI capability growth, but its design isn’t up to the task, argues Nathan Witkin
- Are Short AI Timelines Really Higher-Leverage?
- How well did forecasters predict 2025 AI progress?Mostly right about benchmarks, mixed results on real-world impacts
- AI predictions for 2026But first, scoring my predictions for 2025
- Discarding the Shaft-and-Belt Model of Software DevelopmentAI enables – and benefits from – a move from mega-projects to artisanal solutions
- Why no one can agree on what AI will do to jobsWill AI repeat history — or break it?
- An FAQ on Reinforcement Learning EnvironmentsWe interviewed 18 people across RL environment startups, neolabs, and frontier labs about the state of the field and where it's headed.
- What Happens When Superhuman AIs Compete for Control?...
- Claude Code and What Comes NextWith the right tools, AI can accomplish impressive things
- ML research directions for preventing catastrophic data poisoning
- What sort of post-superintelligence society should we aim for?The case for ‘viatopia’
- Self-sufficient AINo, we don't "have AGI already." But in any case, we should articulate clearer milestones.
- Software Too Cheap to MeterPersonalized solutions may replace one-size-fits-all applications
- Claude Code is about so much more than codingIt’s a general-purpose AI agent. And it’s already a pretty good knowledge worker
- Nine AI predictions for 2026Will the bubble pop? Will AI dominate the midterms? Will Sam Altman snap?
- America's chip export controls are workingDon't be fooled by the people who say we should sell China everything we've got.
- Recent LLMs can do 2-hop and 3-hop latent (no CoT) reasoning on natural factsRecent AIs are much better at chaining together knowledge in a single forward pass
2025
December 2025
- AI Futures Model: Dec 2025 UpdateWe've significantly improved our model of AI timelines and takeoff speeds!
- The AI copyright question has no easy answersNobody can agree on how copyright should apply to AI — or if copyright is fit for purpose at all.
- How far can decentralized training over the internet scale?Decentralized AI training is growing fast, and while it likely won’t catch up to frontier models this decade, even narrowing the gap could have major implications for AI policy.
- Measuring no CoT math time horizon (single forward pass)Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon
- The worst (and funniest) AI takes of 2025feat. Mark Zuckerberg, Gary Marcus, and ... ourselves.
- Chatting with the CorporationWhen “I” stops meaning the bot and starts meaning the org
- The 2025 Transformer Gift GuideThe best presents for the AI-obsessed people in your life.
- Why benchmarking is hardRunning benchmarks involves many moving parts, each of which can influence the final score. The two most impactful components are scaffolds and API providers.
- How the Catholic Church thinks about superintelligenceOpinion: Paolo Benanti, AI advisor to the Vatican and a professor of moral philosophy, argues that there is a moral imperative to ensure that artificial intelligence never becomes humanity’s master
- Can the Pope spur meaningful action on AI?AI safety advocates are hoping an unlikely ally can provide a rallying cry for guardrails and regulation
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceAI can sometimes distribute cognition over many extra tokens
- The changing drivers of LLM adoptionPublic data as well as our original polling suggest LLM adoption is roughly on trend, but the underlying drivers are shifting.
- The Shape of AI: Jaggedness, Bottlenecks and SalientsAnd why Nano Banana Pro is such a big deal
- Presenting the Case That the Future Will Be UnrecognizableWe should prepare for the possibility of profound change
- AI is making dangerous lab work accessible to novices, UK’s AISI findsUK AISI’s first Frontier AI Trends Report finds that AI models are getting better at self-replication, too
- BashArena and Control Setting DesignWe’ve just released BashArena, a new high-stakes control setting we think is a major improvement over the settings we’ve used in the past.
- The GOP consultancy at the heart of the industry’s AI fightHow a network of Targeted Victory alumni are campaigning against AI laws
- Could Space Debris Block Access to Outer Space?New research asks whether deliberately-caused Kessler syndrome could block access to space — and what countermeasures might work.
- Is almost everyone wrong about America’s AI power problem?Why power is less of a bottleneck than you think.
- Checks, Balances, and Power ConcentrationA conversation with Rose Hadshar and Nora Ammann
- The very hard problem of AI consciousnessI spent a weekend at an AI welfare get together — and left with more questions than answers
- So you’ve taken over the worldSuppose that you were running a big AI project that underwent a fast intelligence explosion.
- New York’s governor is trying to turn the RAISE Act into an SB 53 copycatEXCLUSIVE: Gov. Kathy Hochul is proposing to strike the entire text of the RAISE Act, replacing it with verbatim language from SB 53, sources tell Transformer.
- When high scores don’t mean high intelligence: how to build better benchmarksOpinion: Researchers from the Oxford Internet Institute argue that lessons from social sciences can help us better assess the capabilities of AI
- Early US policy priorities for AGINear-term AI policy is confusing, except for these two recommendations
- Why AI reading science fiction could be a problemThe theory that we’re accidentally teaching AI to turn against us
- AI can obviously create new knowledgeBut maybe not new concepts
- Human Dignity: a reviewA manifesto for the age of AI?
- How AI-driven feedback loops could make things very crazy, very fastA primer on the intelligence explosion
- The behavioral selection model for predicting AI motivationsThe basic arguments about AI motivations in one causal graph
- The perils of AI safety’s insularityBy building their own intellectual ecosystem, researchers worried about existential AI risk shed academia's baggage — and, perhaps, some of its strengths
- Another preemption defeat shows the AI industry is fighting a losing battleThe second failed attempt to pass federal preemption of state AI laws could have lasting repercussions for the industry
- How China’s AI diffusion plan could backfireOpinion: Scott Singer argues that the country’s plan to embed AI across all facets of society could create huge growth — and accelerate social unrest
- How important is the model spec if alignment fails?(These are rough research notes.)
- Can AI embrace whistleblowing?As Anthropic prepares to publish its whistleblowing policy, can the industry make the most of protecting those who speak out?
- Thoughts on AI progress (Dec 2025)Why I'm moderately bearish in the short term, and explosively bullish in the long term
- I love AI. Why doesn't everyone?Anti-AI sentiment might or might not be rational, but it certainly relies on a lot of bad arguments.
November 2025
- The environment is a terrible reason to avoid ChatGPTPeople are saying you shouldn’t use ChatGPT due to statistics like:
- A brief guide to the groups protesting over AIThe differences between Stop AI, PauseAI, ControlAI, and more
- Will AI safety become a mass movement?Some AI safety activists think the community should borrow from the climate playbook and build broad public appeal — but not everyone agrees
- SB 53 protects whistleblowers in AI — but asks a lot in returnOpinion: Abra Ganz and Karl Koch argue that whistleblower protections in SB-53 aren’t good enough on the face of it — but how the state chooses to interpret the law could turn that around
- Taking Jaggedness SeriouslyWhy we should expect AI capabilities to keep being extremely uneven, and why that matters
- Will competition over advanced AI lead to war?Fear and Fearon
- Benchmark Scores = General Capability + ClaudinessIs this because skills generalize very well, or because developers are pushing on all benchmarks at once?
- Reflections on The CurveA compilation
- Should the US do a Manhattan Project for AGI?Such a Project is neither inevitable nor a good idea
- Exclusive: Here's the draft Trump executive order on AI preemptionThe EO would establish an “AI Litigation Task Force" to challenge state AI laws
- How profits can drive AI safetyOpinion: Geoff Ralston argues that the safety and security of AI doesn’t need to be at odds with profit and progress
- Hyperproductivity: The Next Stage of AI?A glimpse at an astonishing, exhilarating, exhausting new style of work
- Three Years from GPT-3 to Gemini 3From chatbots to agents
- Why pressure on AI child safety could also address frontier risksKeeping kids safe is a priority for legislators globally — and might increase attention on other risks, too
- RL is even more information inefficient than you thoughtAnd implications for RLVR progress
- Empire of AI is wildly misleading about AI water useAnd the media environment that didn't catch this is getting this issue wrong
- AI Ran Its First Autonomous CyberattackChinese hackers used AI and changed the economics of cyberattacks
- The software intelligence explosion debate needs experimentsThe existing debate rests on data and assumptions that are shakier than most people realize. To make progress, we need better evidence, and experiments are the best way to get it on the margin.
- Will AI systems drift into misalignment?A reason alignment could be hard
- Claude can identify its ‘intrusive thoughts’“I’m experiencing something that feels like an intrusive thought,” Claude said in a recent experiment
- A short summary of my argument that using ChatGPT isn't bad for the environmentTo share with anyone still worried
- Doing AI safety policy when governments aren’t interestedOpinion: Jess Whittlestone argues that there are still ways to keep AI safety policy on the table even when governments don’t prioritize it
- Giving your AI a Job InterviewAs AI advice becomes more important, we are going to need to get better at assessing it
- A Coalition For The FutureIf we can keep it
- AI doesn’t need to be general to be dangerousThere’s more to AI safety than the AGI debate
- The lump of cognition fallacyThe extended mind as the advance of civilization
- AI and Suicide[Warning: Extensive discussion of suicide.
- It's much easier to hold computers accountable than it is to hold humans accountableComputers live in a totalitarian surveillance state
- Data centers and low social trustWhy I was so compelled to post a lot about data centers
- History suggests the AI backlash will failFrom Venetian monks to 19th century textile workers, those opposing new technology have rarely come out on top
- Sora is here. The window to save visual truth is closingOpinion: Sam Gregory argues that generative video is undermining the notion of a shared reality, and that we need to act before it’s lost forever
- A Project Is Not a Bundle of TasksCurrent AIs struggle to create a whole that exceeds the sum of its parts
- An armchair diagnosis of the chatbot moral panicIt's... status!
- OSWorld — AI computer use capabilitiesTasks are simple, many don't require GUIs, and success often hinges on interpreting ambiguous instructions. The benchmark is also not stable over time.
- Why we need to think about taxing AILarge-scale AI-driven unemployment could hit government spending without innovative changes to the tax system
- What's up with Anthropic predicting AGI by early 2027?I operationalize Anthropic's prediction of "powerful AI" and explain why I'm skeptical
October 2025
- Japan’s unusual approach to AI policyThe country’s penalty-free AI legislation relies on social pressure and voluntary compliance — and experts say it could work elsewhere
- Sonnet 4.5's eval gaming seriously undermines alignment evalsAnd this seems caused by training on alignment evals.
- AI is probably not a bubbleAI companies have revenue, demand, and paths to immense value
- Audits, not essays: How to win trust for enterprise AIOpinion: Alexandru Voica argues that application-layer AI companies are best off opening themselves up to rigorous testing rather than opining on AI safety
- What I Saw Around The CurveNotes from the near-future of AI
- What you need to know about the OpenAI restructureNegotiations have seen safeguards and concessions built in, but questions remain about how effective they’ll be, and how fair the deal is
- AI and Folk Cartesianism - Part 2: Problems for CartesianismThe theater is empty, the screen is in another room
- Scenario Scrutiny for AI PolicyA call for concrete stress-testing of AI policy proposals
- Solving AI’s power problem with decentralized trainingConventional wisdom in AI is that large-scale pretraining needs to happen in massive contiguous datacenter campuses. But is this true?
- Where have the really big AI models gone?Why the race to scale up pretraining isn’t over
- AI and Folk Cartesianism - Part 1: Defining the ProblemWhy I'm scared of linear algebra
- Exclusive: UK AISI hires ex-GCHQ AI chief as interim directorAdam Beaumont is replacing Oliver Ilott as head of the UK's AI Security Institute
- How MAGA learned to love AI safetyThe AI industry is engaging in one of “ the most blasphemous endeavors,” a leading MAGA figure told us
- Should AI Developers Remove Discussion of AI Misalignment from AI Training Data?There is some concern that training AI systems on content predicting AI misalignment will hyperstition AI systems into misalignment.
- Is 90% of code at Anthropic being written by AIs?I'm skeptical that Dario's prediction of AIs writing 90% of code in 3-6 months has come true
- Meghan Markle, Steve Bannon and Pope’s AI advisor call for superintelligence ban30% of the American public think superhuman AI should never be developed, too
- Should we worry about AI's circular deals?AI companies are borrowing more money to invest more in AI.
- Thoughts on the AI buildoutFab CapEx overhang, 1 GW a week, China privileged in long timelines, and much else
- How an AI company CEO could quietly take over the world...
- AI cyberrisk might be a bit overhyped — for now at leastExperts say key factors currently limit the risk of catastrophic harm from AI-enabled cyberattacks — as far as we know
- An Opinionated Guide to Using AI Right NowWhat AI to use in late 2025
- Requests for journalists covering AI and the environmentAlways aim to give your reader a full picture of the environmental issues a community and the world is facing
- Less than 70% of FrontierMath is within reach for today’s models57% of problems have been solved at least once
- Reducing risk from scheming by studying trained-in scheming behaviorCan we study scheming by studying AIs trained to act like schemers?
- What happens when the AI bubble bursts?The world is prepping for an AI crash. History points to what that might look like
- A few meta points on my posts on AI and the environmentSome quick notes on what I'm doing
- AI is advancing far faster than our annual report can trackOpinion: Yoshua Bengio, Stephen Clare and Carina Prunkl run through the rapid developments that necessitated an early update to their International AI Safety report
- OpenAI is projecting unprecedented revenue growthNo company has gone from $10B to $100B as fast as OpenAI projects to do
- Bootstrapping to ViatopiaHow short periods of reflection could extend into longer ones, bootstrapping our way to a positive outcome
- AI and synthetic DNA could be a lethal combinationStronger gene-synthesis screening is vital to closing off AI’s ability to enable man-made pandemics
- Gemini 2.5 Deep Think on FrontierMathWe evaluated Gemini 2.5 Deep Think manually on FrontierMath as there is no API. The results: a new record!
- Slop implies capabilityIf the current paradigm can produce convincing, satisfying slop, it can also be useful
- The AI water issue is fakeOn the national, local, and personal level
- Data centers & electricity - part 1: as of 2025 they haven't raised national pricesYet, anyway, but they raised local costs in specific places like NOVA
- Iterated Development and Study of Schemers (IDSS)A strategy for handling scheming
- An intense battle over the RAISE Act is entering its final stretchNew York State's AI bill is more ambitious than California’s SB 53 — and is facing opposition from Andreessen Horowitz and other tech groups
- The RAISE Act can stop the AI industry’s race to the bottomOpinion: Assembly Member Alex Bores argues that regulation can prevent market pressure from encouraging the release of dangerous AI models, without harming innovation.
- The Thinking Machines Tinker API is good news for AI control and securityIt's a promising design for reducing model access inside AI companies.
- Plans A, B, C, and D for misalignment riskDifferent plans for different levels of political will
- What a data center isA building-sized computer, and the most efficient type of building we've ever created
- What the GAIN AI Act could mean for chip exportsA battle is simmering over proposals to make US chipmakers sell to domestic customers before exporting to countries such as China
- How many digital workers could OpenAI deploy?OpenAI has the inference compute to deploy millions of digital workers, but only on a narrow set of tasks – for now.
- Britain’s new AI minister actually ‘gets’ AIKanishka Narayan is excited about AI opportunities, but takes the risks seriously too
- AI models are getting really good at things you do at workA new OpenAI benchmark tests AI models on things people actually do in their jobs — and finds that Claude is about as good as a human for government work
- AI is persuasive, but that’s not the real problem for democracyOpinion: Felix M Simon argues that AI is unlikely to significantly shape election results in the near future, but warns that it could damage democracy through a steady erosion of institutional trust.
September 2025
- Claude Sonnet 4.5 knows when it’s being testedAnthropic's new model appears to use "eval awareness" to be on its best behavior
- When AI starts writing itselfWhy automating AI R&D could be the most dangerous milestone yet
- Real AI Agents and Real WorkThe race between human-centered work and infinite PowerPoints
- More Perfect Union videos are wildly deceptive on data center water useThey're misinforming their huge audience and making the data center debate much worse
- Why GPT-5 used less training compute than GPT-4.5 (but GPT-6 probably won’t)OpenAI focused on scaling post-training on a smaller model
- OpenAI, NVIDIA, and Oracle: Breaking Down $100B Bets on AGIHow vendor financing turns the S&P 500 into a giant AGI bet
- How the UK can seize on Trump’s immigration mistakesOpinion: The H-1B visa changes present a generational opportunity for the UK to scoop up AI talent, Julia Willemyns argues
- Human Drivers Will Kill 11 People While You Read ThisLet’s not obstruct the technology that could have saved them
- Insurance might be the key to making AI secureOpinion: Cristian Trout, Rajiv Dattani and Rune Kvist argue that insurance can help reward responsible development of AI.
- No, ChatGPT isn’t ‘making us stupid’— but there’s still reason to worry
- Notes on fatalities from AI takeoverLarge fractions of people will die, but literal human extinction seems unlikely
- Focus transparency on risk reports, not safety casesTransparency about just safety cases would have bad epistemic effects
- Nobel laureates and AI developers call for ‘red lines’ on AIExperts are calling for international agreements as the United Nations meets, but they face an uphill battle turning words into action
- The world's first frontier AI regulation is surprisingly thoughtful: the EU's Code of PracticeOnly the US can make us ready for AGI, but Europe just made us readier.
- Prospects for studying actual schemersStudying actual schemers seems promising but tricky
- The huge potential implications of long-context inferenceContinual learning, scaling RL, and research feedback loops
- If We Build AI Superintelligence, Do We All Die?If you're not at least a little doomy about AI, you're not paying attention
- Would democracy survive an AGI-supercharged economy?AGI could lead to growth of 30% — and double digit unemployment. It’s not clear that democratic institutions would be able to survive
- Can open-weight models ever be safe?Opinion: Bengüsu Özcan, Alex Petropoulos and Max Reddel argue that technical safeguards, societal preparedness, and new standards could make open-weight models safer
- What training data should developers filter to reduce risk from misaligned AI?An initial narrow proposal
- AI scaling & scientific R&D by 2030What will AI look like by 2030 if current trends hold? What does it mean for AI capabilities in scientific R&D?
- Book Review: 'If Anyone Builds It, Everyone Dies'Eliezer Yudkowsky and Nate Soares’ new book should be an AI wakeup call — shame it’s such a chore to read
- Hiring struggles are plaguing the EU AI OfficeKey leadership roles, including a head of the AI Office safety unit, have yet to be hired
- On Working with WizardsVerifying magic on the jagged frontier
- Three challenges facing compute-based AI policies“Training compute” is constantly evolving, and compute-based AI policies must adapt to remain relevant
- We’re In the Windows 95 Era of AI Agent SecurityHeaded for widespread usage, and insecure by design
- Why AI evals need to reflect the real worldOpinion: Rumman Chowdhury and Mala Kumar argue that we need better AI evaluations — and the infrastructure and investment to do them
- A guide to understanding AI as normal technologyAnd a big change for this newsletter
- Chip location verification is the new export control battlegroundThe Chip Security Act proposes a way to tackle chip smuggling. Semiconductor companies don’t seem to like it.
- We’re getting the argument about AI's environmental impact all wrongIndividual ChatGPT queries are a rounding error — we need to think about the future
- AIs will greatly change engineering in AI companies well before AGIAIs that speed up engineering by 2x wouldn't accelerate AI progress that much
- What's the full "hidden" climate cost of a ChatGPT prompt?It's so tiny, and if we include "all the hidden costs" it might actually be a negative number
- Do data centers only seem bad for the climate because we can see them?Most of your climate impacts are invisible
- California's latest AI safety bill might stand a chanceSB 53 is entering the home stretch despite industry lobbying. Will it make it over the line?
- Compute scaling will slow down due to increasing lead timesA heavily underappreciated dynamic when thinking about AI timelines.
- Explainer: How AI Chips Are MadeIt's complicated
- GPT-5: The Case of the Missing AgentProgress Everywhere Except In The Real World
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me broAbove trend progress due to a rapid increase in RL env quality is unlikely
- Jevons' Paradox is good sometimesIt might be good news for AI and the climate
- Are AI scheming evaluations broken?Doubts have been raised about one of the key ways we tell if AI will misbehave. Is it time for a new approach?
- Compute is a strategic resourceComputing power still determines who wins the AI race
August 2025
- AI and jobs, againSome top economists claim AI is now destroying jobs for a subset of Americans. Are they right?
- Mapping the AI & environment debateAnd clarifying what I believe
- The core simple reason I think AI is valuableWhy I don't think every Watt-hour or liter we spend on AI is wasted
- Mass IntelligenceFrom GPT-5 to nano banana: everyone is getting access to powerful AI
- Attaching requirements to model releases has serious downsides (relative to a different deadline for these requirements)System cards are established but other approaches seem importantly better
- AI embraces crypto’s dirty politicsA new super PAC network looks set to spend millions to influence AI regulation
- An example of what I consider a misleading article about AI and the environmentJust show the numbers!
- Notes on cooperating with unaligned AIsMore thoughts on making deals with schemers
- Why future AI agents will be trained to work togetherMany multi-agent setups are based on fancy prompts, but this is unlikely to persist
- A roundup of AI psychosis storiesConcerning anecdotes keep piling up
- Being honest with AIsWhen and why we should refrain from lying
- Could one country outgrow the rest of the world?When countries grow at the same exponential rate, they maintain their relative sizes. But after we develop AGI, there may be a period of superexponential growth, with growth becoming faster and faster over time. If this superexponential growth lasts for long enough, the leader could pull further…
- My AGI timeline updates from GPT-5 (and 2025 so far)AGI before 2029 now seems substantially less likely
- Anthropic’s piracy could make its copyright battle existentialAnthropic could have to pay billions in damages — and others may follow
- How to make the future better (other than by reducing extinction risk)A summary of a new essay.
- 35 Thoughts About AGI and 1 About GPT-5Such As: How do Teenagers Learn to Drive 10,000x Faster Than Waymo?
- Contra the UK government, please don't delete your old photos and emails to save waterYou'd need to delete hundreds of billions of emails to save as much water as fixing your toilet
- Do we understand how neural networks work?Yes and no, but mostly no.
- AGI: Probably Not 2027AI 2027 is a web site that might be described as a paper, manifesto or thesis.
- Donald Trump is making Chinese AI great againGiving China access to advanced chips for a $2bn payoff is a bad deal for US security
- Projecting AI Training Power DemandWhat happens if trends continue?
- The trajectory of the future could soon get set in stoneA summary of our new paper, "Persistent Path-Dependence."
- Four places where you can put LLM monitoringTo wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems
- GPT-5: a small step for intelligence, a giant leap for normal peopleGPT-5 focuses on where the money is - everyday users, not AI elites
- Will morally motivated actors steer us towards a near-best future?A summary of our new paper, “Convergence and Compromise”.
- GPT-5: It Just Does StuffPutting the AI in Charge
- How Quick and Big Would a Software Intelligence Explosion Be?AI systems may soon fully automate AI R&D.
- Why AI's IMO gold medal is less informative than you thinkThe problems gave AI only a slim chance to show new capabilities
- Is eutopia the default outcome post-AGI?In the new paper “No Easy Eutopia”, we argue that reaching a truly great future will be very hard.
- Z.ai and Huawei aren't defeating US export controlsLoosening restrictions now would surrender America's technological advantage
- Does Trump’s AI Action Plan have what it takes to win?Ten takes on the AI Action Plan
- Should we aim for flourishing over mere survival?The Better Futures series.
- Quantifying the algorithmic improvement from reasoning modelsReasoning models were as big of an improvement as the Transformer, at least on some benchmarks
July 2025
- UK launches £15 million AI alignment project“Alignment is one of the most urgent technical challenges of our time,” the government said.
- Should we update against seeing relatively fast AI progress in 2025 and 2026?Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.
- The Bitter Lesson versus The Garbage CanDoes process matter? We are about to find out.
- Why China isn’t about to leap ahead of the West on computeChinese hardware is closing the gap, but major bottlenecks remain
- Grok 4’s math capabilitiesAn assessment of math capabilities beyond headline numbers
- AI As Profoundly Abnormal Technology....
- Official White House policy: AI is a big dealThe AI vibe shift continues. Action is next.
- Why is Hugging Face hosting tools to make deepfake porn of teenage celebrities?The “ethical” AI company is still hosting models designed to make nonconsensual sexual content, despite clear breaches of its policies
- Personalized AI is rerunning the worst part of social media's playbookThe incentives, risks, and complications of AI that knows you
- We aren't worried about misalignment as self-fulfilling prophecy...
- Why it's hard to make settings for high-stakes control researchIt's like making challenging evals, but more constrained
- After the ChatGPT Moment: Measuring AI’s AdoptionHow quickly has AI been diffusing through the economy?
- Could AI slow science?Confronting the production-progress paradox
- What Makes AI "Generative"?To "generate" is to create, either from nothing (The Book of Genesis), or from very different or relatively inactive materials (an electrical generator, generating offspring).
- Recent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- AI is the most rapidly adopted technology in history7 charts of real world AI deployment
- Not So Fast: AI Coding Tools Can Actually Reduce ProductivityStudy Shows That Even Experienced Developers Dramatically Overestimate Gains
- Can we safely deploy AGI if we can't stop MechaHitler?We need to see this as a canary in the coal mine
- What will the IMO tell us about AI math capabilities?Most discussion about AI and the IMO focuses on gold medals, but that's not the thing to pay most attention to.
- What's worse, spies or schemers?And what if you have both at once?
- Against "Brain Damage"AI can help, or hurt, our thinking
- How much novel security-critical infrastructure do you need during the singularity?And what does this mean for AI control?
- Two proposed projects on abstract analogies for schemingWe should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.
- How big could an “AI Manhattan Project” get?An AI Manhattan Project could accelerate compute scaling by two years
- There are two fundamentally different constraints on schemers"They need to act aligned" often isn't precise enough
- On The Platonic Representation HypothesisEmpirically there is one and only one correct understanding of this, and every other, post.
June 2025
- Unresolved debates about the future of AIHow far the current paradigm can go, AI improving AI, and whether thinking of AI as a tool will keep making sense
- What you can do about AI 2027How to steer toward a positive AGI future
- Congress has started taking AGI more seriouslyThe AGI vibes, they are a-shiftin'
- Jankily controlling superintelligenceHow much time can control buy us during the intelligence explosion?
- The Industrial ExplosionOnce AI can automate human labour, positive feedback loops could lead to a rapid increase in physical capabilities as well as cognitive ones.
- AI can be bad without being uselessA common way conversations get tripped up
- How not to lose your job to AIThe skills AI will make more valuable (and how to learn them)
- My "Are you presuming most people are stupid?" testFor AI criticism and everything else
- What does 10x-ing effective compute get you?Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?
- Comparing risk from internally-deployed AI to insider and outsider threats from humansAnd why I think insider threat from AI combines the hard parts of both problems.
- Using AI Right Now: A Quick GuideWhich AIs to use, and how to use them
- AI and explosive growth reduxTwo updates from our integrated assessment model of AI automation
- Making deals with early schemers...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.
- Prefix cache untrusted monitors: a method to apply after you catch your AITraining the policy to not do egregious bad actions we detect has downsides and we might be able to do better
- AI safety techniques leveraging distillationDistillation is cheap; how can we use it to improve safety?
- Computing is efficientThe reason using ChatGPT isn't bad for the environment is that it's a computer program
- Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?Anthropic has triggered "AI Safety Level 3" protections - but do their evaluations support this decision?
- What does SWE-bench Verified actually measure?It is one of the best tests of AI coding, but limited by its focus on simple bug fixes in familiar repositories
- When does training a model change its goals?Can a scheming AI's goals really stay unchanged through training?
- Beyond benchmark scores: Analyzing o3-mini’s mathematical reasoningWhy o3-mini is a "vibes-based inductive reasoner"
- The Biggest Statistic About AI Water Use Is A LieHow did it become the main story?
- AI History in QuotesEach of these presents the clearest, earliest, or most-cited statement of a specific idea in AI.
- Building supercomputers for autocrats probably isn’t good for democracy, actuallyThe shameless spin around OpenAI's Stargate UAE
- Anthropic C.E.O.: Don’t Let A.I. Companies off the Hook“But a 10-year moratorium is far too blunt an instrument. A.I. is advancing too head-spinningly fast. I believe that these systems could change the world, fundamentally, within two years; in 10 years, all bets are off. Without a clear plan for a federal response, a moratorium would give us the worst of both worlds — no ability for states to act, and no national policy as a backstop.”
- What Is AI?This is an important question because AI is, currently, important.
- Why I don’t think AGI is right around the cornerContinual learning is a huge bottleneck
- The recent history of AI in 32 ottersThree years of progress as shown by marine mammals
May 2025
- Lab vs. Life: Dissecting “AI as Normal Technology”What if AGI Can't Be Developed in an Ivory Tower?
- Give AIs a stake in the futurePart 1 of Classical Liberal AGI
- GPQA Diamond: what’s left?Reports of Its Death Are Somewhat Exaggerated
- Human Takeover Might be Worse than AI TakeoverSome are concerned that future AI systems could take over the world.
- The case for countermeasures to memetic spread of misaligned valuesDefending against alignment problems that might come with long-term memory
- All the ways I want the AI debate to be betterGround truths, useful rules, ideas I'd like taken seriously, and a rant
- Is AI already superhuman on FrontierMath?How do humans and AIs compare on FrontierMath? We ran a competition at MIT to put this to the test.
- Making AI Work: Leadership, Lab, and CrowdA formula for AI in companies
- When Decades Become Days: Dissecting AI 2027The Four Requirements for a Short AI Timeline
- Misaligned AI is no longer just theoryA host of new evidence shows that misalignment is possible — but it's unclear whether harm will follow
- Reactions to MIT Technology Review's report on AI and the environmentInteresting useful facts, and some framing I think is super misleading
- Book Review: ‘The Optimist’ and ‘Empire of AI’Two new books try to unmask Sam Altman and OpenAI, to varying degrees of success
- Slow corporations as an intuition pump for AI R&D automationIf slower employees would be much worse wouldn't automated faster ones be much better?
- Make The Prompt PublicThe right to know what an AI is doing
- How fast can algorithms advance capabilities?How much do the best algorithmic innovations depend on compute?
- Things I got wrong in my ChatGPT/environment postsLogging the errors
- We need to know what’s happening with AIWithout transparency, society is flying blind toward potentially catastrophic AI capabilities
- AIs at the current capability level may be important for future safety workSome reasons why relatively weak AIs might still be important when we have very powerful AIs
- In search of a dynamist vision for safe superhuman AIEmbracing creativity, risk-taking, and competition while staying clear-eyed about risks
- Why goofy AI art almost never seems wasteful to meEnvironmentalism is important, but shouldn't be used to squash pluralism
- How far can reasoning models scale?Available evidence suggests that rapid growth in reasoning training can continue for a year or so.
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration HackingA new analysis of the risk of AIs intentionally performing poorly.
- Replies to criticisms of my posts on ChatGPT & the environmentWhy I used 3 Wh, and why I don't say "Every Watt-hour matters"
- Training-time schemers vs behavioral schemersClarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.
- Where’s my ten minute AGI?Why don’t AIs automate more real-world tasks if they can handle 1-hour ones? Here are at least three fundamental reasons.
- What's going on with AI progress and trends? (As of 5/2025)My views on what's driving AI progress and where it's headed.
- The crucibleHow I think about the situation with AI
- AGI is not a milestoneThere is no capability threshold that will lead to sudden impacts
- Making sense of OpenAI's modelsPlus: GPT-5's secret identity
- Personality and PersuasionLearning from Sycophants
April 2025
- How can we solve diffuse threats like research sabotage with AI control?Preventing research sabotage will require techniques very different from the original control paper.
- 7+ tractable directions in AI controlA list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute
- First, They Came for the Software Engineers…What 25 interviews tell us about AI's messy impact on tech work
- Should you quit your job – and work on risks from AI?In five years, we could have AI systems capable of accelerating science and automating skilled jobs.
- Using ChatGPT is not bad for the environment - a cheat sheetThe numbers clearly show that discouraging people from using chatbots is a pointless distraction for the climate movement
- The case for multi-decade AI timelinesIn this Gradient Updates weekly issue, Ege discusses the case for multi-decade AI timelines.
- Clarifying AI R&D threat models(There are a few)
- How training-gamers might function (and win)A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.
- Finally, A Way to Measure AI Progress; Everyone's Misreading ItThat METR Study Doesn’t Say "AGI in 5 Years"
- The Urgency of Interpretability“People outside the field are often surprised and alarmed to learn that we do not understand how our own AI creations work. They are right to be concerned: this lack of understanding is essentially unprecedented in the history of technology.”
- Why America WinsIn a word: compute
- 2 big questions for AI progress in 2025-2026On how good AI might—or might not—get at tasks beyond math & coding
- AI 2027: Media, Reactions, CriticismWe recognize our supporters and respond to our critics
- ChatGPT on the lives of factory farmed pigsAn example of deep research
- Forecaster reacts: METR's bombshell paper about AI accelerationNew data supports an exponential AI curve, but lots of uncertainty remains
- Questions about the Future of AIConsiderations about economics, history, training, deployment, investment, and more
- On Jagged AGI: o3, Gemini 2.5, and everything afterNew models and new thresholds
- Handling schemers if shutdown is not an optionWhat if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?
- Training AGI in Secret would be Unsafe and UnethicalBad for loss of control risks, bad for concentration of power risks
- AI-enabled coups: how a small group could use AI to seize powerThe development of AI that is more broadly capable than humans will create a new and serious threat: AI-enabled coups. An AI-enabled coup is most likely to be staged by leaders of frontier AI projects, heads of state, and military officials; and could occur even in established democracies.
- Ctrl-Z: Controlling AI Agents via ResamplingA new paper on AI control for agents.
- AI as Normal TechnologyA new paper that we will expand into our next book
- To be legible, evidence of misalignment probably has to be behavioralEvidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.
- Beyond The Last HorizonWhat are time horizons, and how do we use them in our forecast?
- Why do misalignment risks increase as AIs get more capable?A breakdown of how higher capabilities increase risk
- Disempowerment spiralsNotes on a likely mechanism for existential catastrophe
- An overview of areas of control workWhat are all the research and implementation areas helpful for control?
- Shortening AGI timelines: a review of expert forecastsAs a non-expert, it would be great if there were experts who could tell us when we should expect artificial general intelligence (AGI) to arrive.
- How I use AILLMs in everyday life and work
- An overview of control measuresWhat methods can we use to ensure control?
- Will we have AGI by 2030?In recent months, the CEOs of leading AI companies have grown increasingly confident about rapid progress:
- Nonproliferation is the wrong approach to AI misuseMaking the most of “adaptation buffers” is a more realistic and less authoritarian strategy
- Notes on countermeasures for exploration hacking (aka sandbagging)How can we prevent AIs from intentionally underperforming on our metrics?
- Our first project: AI 2027What superintelligence looks like
- The core challenge of AI alignment is “steerability”"Steered to where" is a different question from "steerable at all"
- Mutual sabotage of AI probably won’t workAI deterrence isn’t like nuclear deterrence
- The AI Adoption Gap: Preparing the US Government for Advanced AIAdvanced AI could unlock an era of enlightened and competent government action. But without smart, active investment, we’ll squander that opportunity and barrel blindly into danger.
- "Long" timelines to advanced AI have gotten crazy shortThe prospect of reaching human-level AI in the 2030s should be jarring
March 2025
- No elephants: Breakthroughs in image generationWhen Language Models Learn to See and Create
- Notes on handling non-concentrated failures with AI control: high level methods and different regimesWhat are the methods and issues when failures occur diffusely over many actions?
- The real reason AI benchmarks haven’t reflected economic impactsThe real reason that AI benchmarks haven’t reflected real-world impacts historically is that they weren’t optimized for this, not because of fundamental limitations – but this might be changing.
- Will the Need to Retrain AI Models from Scratch Block a Software Intelligence Explosion?Once AI fully automates AI R&D, there might be a period of fast and accelerating software progress – a software intelligence explosion (SIE).
- Knowledge, Reasoning, and SuperintelligenceWhat makes people good at solving novel problems?
- Will AI R&D Automation Cause a Software Intelligence Explosion?Empirical evidence suggests that, if AI automates AI research, feedback loops could overcome diminishing returns, significantly accelerating AI progress.
- How to make AI go well: a summaryI’m writing a new guide to careers to help AGI go well. Here's a summary of the key messages as they stand.
- The Cybernetic TeammateHaving an AI on your team can increase performance, provide expertise, and improve your experience
- Most AI value will come from broad automation, not from R&DAI's biggest impact will come from broad labor automation—not R&D—driving economic growth through scale, not scientific breakthroughs.
- Should There Be Just One Western AGI Project?There have been recent discussions of centralizing western AGI development, for instance through a Manhattan Project for AI.
- A defense of AI artIt's not just slop, it's not stolen, it's not bad for the environment, and we should want art to be easy to make
- The quest to build better defenses for AI risks'Societal resilience' measures might offer some protection to the proliferation of dangerous AI capabilities
- The most important graph in AI right now: time horizonTo understand how close we are to transformative AI, here’s the metric I find most interesting right now: how long are the tasks AI can do?
- AI Tools for Existential SecurityRapid AI progress is the greatest driver of existential risk in the world today. But — if handled correctly — it could also empower humanity to face these challenges.
- Prioritizing threats for AI controlWhat are the main threats and how should we prioritize them?
- AI Tools for Existential Security(From a paper coauthored with Lizka Vaintrob.)
- Three Types of Intelligence ExplosionOnce AI systems can design and build even more capable AI systems, we could see an intelligence explosion, where AI capabilities rapidly increase to well past human performance.
- Intelsat as a Model for International AGI GovernanceIf there is an international project to build artificial general intelligence (“AGI”), how should it be designed?
- Being responsible with Chinese AI hypeLessons from Manus and the dangers of uncritical tech narratives
- Preparing for the Intelligence ExplosionImagine all the scientific, intellectual and technological developments that you would expect to see by the year 2125, if technological progress continued over the next century at roughly the same rate that it did over the last century.
- Speaking things into existenceExpertise in a vibe-filled world of work
- What AI can currently do is not the storyForecasting AI progress requires more than extrapolating current capabilities; understanding fundamental task difficulty is key to predicting future breakthroughs.
- AGI by 2030? What Policy Leaders, Tech Leaders, and Pokémon sayThree podcasts and one Pokémon reset show where AI is going
February 2025
- The promise of reasoning modelsReasoning models are perhaps best understood as part of a broader, longer-term trend in which AI systems incrementally take on new tasks they were previously incapable of handling.
- AI coding tools are quietly reshaping software developmentThey’re making an economic impact, despite not being very good yet
- If AGI Means Everything People Do... What is it That People Do?And Why Are Today’s "PhD" AIs So Hard To Apply To Everyday Tasks?
- AI security is important practice for when stakes go upWhy today's safeguards matter for future AI capabilities
- A new generation of AIs: Claude 3.7 and Grok 3Yes, AI suddenly got better... again
- AI progress is about to speed upAI progress is accelerating, with next-gen models surpassing GPT-4 in compute power, driving major leaps in reasoning, coding, and math capabilities.
- How might we safely pass the buck to AI?Developing AI employees that are safer than human ones
- We're Finding Out What Humans are Bad AtAI Advances Fastest When We Find Unnatural Ways of Doing Things
- AI excels at code competitions, struggles with real workWhat CodeForces rankings reveal about AI capabilities
- Algorithmic progress likely spurs more spending on compute, not lessAlgorithmic progress in AI may not reduce compute spending—instead, it could drive higher investment as efficiency unlocks new opportunities.
- Decentralized training isn't a policy nightmare — yetGovernments should still be able to keep track of who's training frontier models — though a shift to reinforcement learning could make that harder
- Ten Takes on the Paris AI Action SummitFrance wants to race, but isn't looking at the road ahead
- Teaching AI to reason: this year's most important storyMost people think of AI as a pattern-matching chatbot – good at writing emails, terrible at real thinking.
- Ideas from philosophy I use to think about AIPhysicalism, functionalism, and lots of Quine
- The embarrassing failure of the Paris AI SummitExperts are sounding the alarm — but governments simply won’t listen
- The promises and perils of voluntary commitments for AI safetyIt's a better place to start than you might think, but stronger measures will soon be necessary
- Gary Marcus says AI can't do things it can already doThe problem of criticising AI using outdated models
- Leaked: this is the AI Action Summit statementThe statement, set to be signed by countries next week, is a 'wasted opportunity', experts say — and looks unlikely to be signed by US officials
- How much energy does ChatGPT use?This Gradient Updates issue explores how much energy ChatGPT uses per query, revealing it's 10x less than common estimates.
- The End of Search, The Beginning of ResearchThe first narrow agents are here
- Ten Takes on DeepSeekNo, it is not a $6M model nor a failure of US export controls
January 2025
- What fully automated firms will look likeEveryone is sleeping on the *collective* advantages AIs will have, which have nothing to do with raw IQ: they can be copied, distilled, merged, scaled, and evolved in ways humans simply can't.
- What went into training DeepSeek-R1?On January 20th, 2025, DeepSeek released their latest open-weights reasoning model, DeepSeek-R1, which is on par with OpenAI’s o1 in benchmark performance.
- China’s DeepSeek Adds a Weird New Data Point to The AI RaceV3 and R1 are Impressive Work, With Many Implications – but not "China Has Caught Up"
- Takeaways from sketching a control safety caseInsights from a long technical paper compressed into a fun little commentary
- Exclusive: Americans overwhelmingly support AI safety mandates, new poll findsThe public would rather ban AI development than have no regulation at all
- On DeepSeek and Export Controls“Export controls serve a vital purpose: keeping democratic nations at the forefront of AI development. To be clear, they’re not a way to duck the competition between the US and China.”
- Planning for Extreme AI RisksAre we ready for this?
- Ten people on the insideA scary scenario that's worth planning for
- Which AI to Use Now: An Updated Opinionated Guide (Updated Again 2/15)Picking your general-purpose AI
- AGI could drive wages below subsistence levelHistorically, many have feared that automation would lead to mass unemployment and lower wages.
- The way we evaluate AI model safety might be about to breakAs systems become more capable, researchers think we need a new type of safety evaluation
- When does capability elicitation bound risk?The assumptions behind and limitations of capability elicitation have been discussed in multiple places (e.g.
- Does Elon still care about AI safety?He’s doing a terrible job of showing it.
- How will we update about scheming?A quantitative description of how I expect to change my mind.
- How has DeepSeek improved the Transformer architecture?DeepSeek has recently released DeepSeek v3, which is currently state-of-the-art in benchmark performance among open-weight models, alongside a technical report describing in some detail the training of the model.
- Thoughts on the conservative assumptions in AI controlWhy are we so friendly to the red team?
- Extending control evaluations to non-scheming threatsBuck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model psychologies are not well described by the standard notion of…
- Using ChatGPT is not bad for the environmentAnd a plea to think seriously about climate change without getting distracted
- How quickly could robots scale up?Some notes on robot economics.
- AI Generated Misinformation is Still a RiskHalf the world's population voted in 2024, relatively unaffected by AI generated deepfakes. Can we conclude that AI generated misinformation isn't the risk we thought it was?
- New AI export controls have leaked. Here’s what you need to know.The new rules are squarely focused on the frontier
- Prophecies of the FloodWhat to make of the statements of the AI labs?
- The economic consequences of automating remote workRecent AI progress has shown great promise in automating cognitive tasks, like those in natural language processing and vision.
- Are We on the Brink of AGI?A Tale of Two Timelines
- Trump Can Keep America’s AI Advantage
- The Important Thing About AGI is the Impact, Not the NameReality Doesn't Care How We Interpret the Words "General Intelligence"
2024
December 2024
- The media needs to start taking AGI seriouslyAn essay for Nieman Lab's 2025 Predictions series
- Moravec’s paradox and its implicationsSince the birth of the field of artificial intelligence in the 20th century, researchers have observed that the difficulty of a task for humans at best weakly correlates with its difficulty for AI systems.
- LLMs Fight With Both Hands Tied Behind Their BackWe haven't yet given them access to knowledge-in-the-world
- The Black Spatula Project: Day FiveOff to a Roaring Start
- How to prepare yourself for AGIWhat if artificial general intelligence (AGI) arrives in just a couple of years, triggering an explosion in science and technology that transforms life as we know it?
- How do mixture-of-experts models compare to dense models in inference?In last week’s Gradient Updates issue, I discussed how we can guess that GPT-4o and Claude 3.5 Sonnet have significantly fewer parameters than GPT-4.
- Measuring whether AIs can statelessly strategize to subvert security measuresThe complement to control evaluations
- What just happenedA transformative month rewrites the capabilities of AI
- Alignment Faking in Large Language ModelsIn our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.
- Is AI progress slowing down?Making sense of recent technology trends and claims
- The Black Spatula ProjectFixing flawed scientific papers, 5000 tokens at a time
- The Future Is Already Here, It’s Just Not Evenly DistributedYou Can Awe Some of the People Some of the Time
- Frontier language models have become much smallerBetween the release of the original Transformer in 2017 and the release of GPT-4, language models at the frontier of capabilities became much larger.
- We Looked at 78 Election Deepfakes. Political Misinformation is not an AI Problem.Technology Isn’t the Problem—or the Solution.
- 15 Times to use AI, and 5 Not toNotes on the Practical Wisdom of AI Use
- What did US export controls mean for China’s AI capabilities?Four days ago, the US government announced new rules around the export of powerful chips and semiconductor manufacturing equipment to China.
- OpenAI's new model tried to avoid being shut downo1 attempted to exfiltrate its weights to avoid being shut down
- The 2024 Transformer Gift GuideThe must-buy items for the AI policy people in your life
November 2024
- Why a US AI "Manhattan Project" could backfire: notes from conversations in ChinaMy recent two weeks in China suggested something surprising about its AI landscape: the biggest bottleneck isn't compute – it's commitment.
- Getting started with AI: Good enough promptingDon't make this hard
- OpenAI's CBRN tests seem unclearOpenAI says o1-preview can't meaningfully help novices make chemical and biological weapons. Their test results don’t clearly establish this.
- Synthetic data is more useful than you thinkFears of ‘model collapse’ are overstated — at least for now
- What Are the Real Questions in AI?It's hard to have a constructive discussion if we're not talking about the same thing
- The Choice TransitionOn the emergence of history’s reins
- Why imperfect adversarial robustness doesn't doom AI controlThere are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.
- New OpenAI emails reveal a long history of mistrustGreg Brockman and Ilya Sutskever had questions about Sam Altman's intentions as early as 2017
- Win/continue/lose scenarios and execute/replace/audit protocolsIn this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.
- Yet another AI safety researcher has left OpenAIRichard Ngo resigned today, saying it has become "harder for me to trust that my work here would benefit the world"
- Meta’s AI ‘safeguards’ are an elaborate fictionMeta cannot prevent misuse, despite what it might pretend
- AI is Racing Forward – on a Very Long RoadProgress Hasn’t Stalled; But AGI Is Not Yet Near
- Does the UK’s liver transplant matching algorithm systematically exclude younger patients?Seemingly minor technical decisions can have life-or-death effects
- Why AI companies are eyeing the Middle EastAnd what the US government is doing about it
- What Trump means for AI safetyA repeal of last year's executive order, for one thing
- A brief history of the automated corporationLooking back from 2041
- The Present Future: AI's Impact Long Before SuperintelligenceYou can start to see the outlines of an AI future, for better and worse
October 2024
- It’s time to take AI welfare seriouslyA new report argues that AI systems could soon deserve moral consideration
- Anthropic has hired an 'AI welfare' researcherKyle Fish joined the company last month to explore whether we might have moral obligations to AI systems
- To Change the World, Set a Bold Target: Moore's Law as Self-Fulfilling ProphecyAll progress depends on the unreasonable forecast
- Abandon compute thresholds at your perilThey are the best tool we have to reduce the burden of AI regulation
- AI safety tax dynamicsEarlier AI capabilities — research and coordination — could help us to navigate later safety problems. And the highest risk probably comes in the era before strong superintelligence.
- When you give a Claude a mouseSome quick impressions of an actual agent
- Is OpenAI being fair to its non-profit?Experts fear OpenAI’s conversion to a for-profit might not adequately compensate its charitable arm
- Thinking Like an AIA little intuition can help
- Safety tax functionsA conceptual exploration
- No, A Bot Didn't Just Make $150M in CryptoBeware the Startling Anecdote
- What would evidence-based AI policy look like?If-then commitments might be the answer
- Learning to Explore: AlphaProof and o1 Show The Path to AI CreativityAn Impressive Start, But Many Hurdles Lie Ahead
- A Policy Agenda for Defensive Acceleration Against AI RisksThe same AI capabilities underlie benefits and harms. How can we assure AI safety without curtailing beneficial applications? Defensive acceleration charts a possible course.
- Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingOne strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).
- How I use AI as a journalistLLMs have become an indispensable part of my workflow
- What if everyone is wrong about what AI does?Both critics and supporters seem to think AI is a "human remover". What if they're both wrong?
- AI Creativity Is A Question of Quality, Not NoveltyIt Actually Matters How Well The Bear Dances
September 2024
- Gavin Newsom has caved to the billionairesSB 1047's veto is a win for tech companies — and a loss for everyone else
- A basic systems architecture for AI agents that do autonomous researchAnd diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.
- How to prevent collusion when using untrusted models to monitor each otherSuppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions.
- Can AI automate computational reproducibility?A new benchmark to measure the impact of AI on improving science
- Lies and deception: Andreessen Horowitz’s SB 1047 campaign is as misleading as it getsThe venture capital giant has repeatedly peddled untruths in an effort to kill the AI regulation bill
- What It’s Like To Solve a Math Olympiad ProblemAnd What This Tells Us About Creativity, Reasoning, And AI
- Scaling: The State of Play in AIA brief intergenerational pause...
- AI Could Break Things; Let's Use It As a Wakeup Call To Make Them StrongerStop Treating AI Policy Tradeoffs As Zero-Sum
- OpenAI's new models 'instrumentally faked alignment'The o1 safety card reveals a range of concerning capabilities, including scheming, reward hacking, and biological weapon creation.
- Something New: On OpenAI's "Strawberry" and ReasoningSolving hard problems in new ways
- AI, centralization, and the One RingConcerns with amassing power in a single AGI project
- AI employees are defying their employers to support SB 1047In speaking out, they’re showing just how important the bill is
August 2024
- Rather Than Arguing About What We Don't Know, Let's Work Together To Find OutA Lesson From The SB 1047 Debate
- Post-apocalyptic educationWhat comes after the Homework Apocalypse
- Would catching your AIs trying to escape convince AI developers to slow down or undeploy?I'm not so sure.
- Dangerous capability tests should be harderWe should spend less time proving that today’s AIs are safe and more time figuring out how to tell if tomorrow’s AIs are dangerous.
- AI companies are pivoting from creating gods to building products. Good.Turning models into products runs into five challenges
- Fields that I reference when thinking about AI takeover preventionIs AI takeover like a nuclear meltdown? A coup? A plane crash?
- Change blindness21 months later
- Is the UAE running an AI influence campaign?Twitter is overrun with bot-like accounts praising the UAE's approach to AI
- Resilience and Adaptation to Advanced AIWhen model safeguards fail, what's our backup plan? AI resilience is crucial for preparing for the widespread diffusion of increasingly powerful AI.
- On speaking to AIVoice changes a lot of things
- Senate Commerce Committee advances lots of AI billsTed Cruz isn't happy about it, though
July 2024
- AI companies are falling short on their promises to the White HouseA year on from making voluntary commitments to the White House, adherence is patchy
- AI existential risk probabilities are too unreliable to inform policyHow speculation gets laundered through pseudo-quantification
- Yes, we still have to workThe automated luxury paradise is still just science fiction.
- The last era of human mistakesAuthor’s remark, four months later: in retrospect I am vaguely dissatisfied with this piece.
- Confronting Impossible FuturesWe shouldn't be certain about what is next, but we should plan for it
- What Kamala Harris means for AI regulationThe vice-president has acknowledged the existential risks of AI and called for legislation to tackle those risks
- Meta-funded group floods Facebook with anti-AI regulation adsThe American Edge Project has spent $150k+ stoking China fears and urging opposition to "anti-innovation" laws
- Decomposing AgencyDesires without capabilities; capabilities without desires
- Tech companies are trying to kill California's AI regulation billAndreessen Horowitz, Y Combinator, and Big Tech are trying to stop SB 1047 from passing
- What Labour means for AIEverything you need to know about Keir Starmer and Peter Kyle’s plans
- Gradually, then Suddenly: Upon the ThresholdSmall improvements can lead to big changes
- New paper: AI agents that matterRethinking AI agent benchmarking and evaluation
- Lawrence Lessig is very worried about freely available AI model weights"Open weights create a unique kind of risk," says the open source pioneer
June 2024
- AI scaling mythsScaling will run out. The question is when.
- Latent Expertise: Everyone is in R&DIdeas come from the edges, not the center
- Ilya Sutskever is betting on very cheap superintelligenceReading between the lines of "Safe Superintelligence Inc".
- Getting 50% (SoTA) on ARC-AGI with GPT-4oYou can just draw more samples
- The only AI certainty is uncertaintyA response to Jack Clark's GPT-2 reflections
- AI takeoff and nuclear warDrivers of all-out war, and strategies to mitigate the risk
- On the future of language modelsMechanistic but zoomed-out takes about the trajectory of the technologies
- What Apple's AI Tells Us: Experimental Models⁴Siri versus the machine god?
- Access to powerful AI might make computer security radically easierAI might be really helpful for reducing security risk.
- A Response to "Situational Awareness"I think the scenario is wrong; let's come up with a plan that works either way
- Doing Stuff with AI: Opinionated Midyear EditionAI systems have gotten more capable and easier to use
- OpenAI employee says he was fired for raising security concerns to boardLeopold Aschenbrenner also said he was interrogated about his team’s “loyalty to the company”
- AI catastrophes and rogue deploymentsIt’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".
- Scientists should use AI as a tool, not an oracleHow AI hype leads to flawed research that fuels more hype
May 2024
- Grounding the Conversation About AIOrganizing The (AI) World’s Disagreements And Making Them Universally Accessible and Constructive
- Wiener: SB 1047 is 'not looking to cover startups'At a town hall event, California Sen. Scott Wiener said the bill is facing pushback from big tech companies
- Sam Altman was 'outright lying to the board', says former board memberIn an interview with TED, Helen Toner said that four OpenAI board members "couldn't believe things that Sam was telling us"
- Four Singularities for ResearchThe rise of AI is creating both crisis and opportunity
- Alliance for the Future director Brian Chau has history of racist, sexist remarksIn a post on AI regulation, Chau falsely claimed George Floyd was a "domestic abuser"
- The most interesting startup idea I've seen recently: AI for epistemicsIf transformative AI might come soon and you want to help that go well, one strategy you might adopt is building something that will improve as AI gets more capable.
- Meet Meta's AI lobbying armyWith 30 lobbyists and seven agencies, the company is primed to push its agenda on Washington
- OpenAI is haemorrhaging safety talentSince the failed Altman ouster, many of the company's most safety-conscious employees have left
- The “ethics vs safety” fight misses the real enemy: Big TechWhile advocates for regulation squabble, Big Tech is pushing for no rules at all
- What OpenAI didA new model opens up new possibilities
- How much AI inference can we do?Suppose you have a bunch of GPUs.
- Superhuman?What does it mean for AI to be better than a human? And how can we tell?
- Preventing model exfiltration with upload limitsUnlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.
- Catching AIs red-handedIf your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.
- Managing catastrophic misuse without robust AIHow could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?
- The case for ensuring that powerful AIs are controlledLabs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.
- Untrusted smart models and trusted dumb modelsThe easiest way to be sure an AI isn't scheming against you is to note that it's too dumb to pull that off. What happens if that's the only way we have to rule out scheming?
- One Conversation is Worth a Thousand Angry TakesMake The Internet Better Using This One Weird Trick
- Freeing the chatbotIntelligence, of a sort, is going to be all around us
April 2024
- AI leaderboards are no longer useful. It's time to switch to Pareto curves.What spending $2,000 can tell us about evaluating AI agents
- AI stocks could crashAnd this could have implications for timelines and AI safety.
- The market expects AI software to create trillions of dollars of value by 2027Does that seem high or low to you?
- What To Expect When You’re Expecting GPT-5Where Will AI Go From Here?
- Innovation through promptingDemocratizing educational technology... and more
- What just happened, what is happening nextThe tasks AI can do well are expanding rapidly
- Tech policy is only frustrating 90% of the timeThat’s what makes it worthwhile
March 2024
- On the necessity of a sinWhy treating AI like a person is the future
- Which AI should I use? Superpowers and the State of PlayAnd then there were three
- I Don't See How Comparative Advantage Applies In a World of Strong AIEconomic models are based on simplifying assumptions that may not hold in an AI-saturated future
- I, Cyborg: Using Co-IntelligenceHow I used AI in my book about AI
- AI safety is not a model propertyTrying to make an AI model that can’t be misused is like trying to make a computer that can’t be used for bad things
- A safe harbor for AI evaluation and red teamingAn argument for legal and technical safe harbors for AI safety and trustworthiness research
- Captain's log: the irreducible weirdness of prompting AIsAlso, we have a prompt library!
February 2024
- On the Societal Impact of Open Foundation ModelsAdding precision to the debate on openness in AI
- Book review: "Power and Progress"In which Daron Acemoglu and Simon Johnson fail to convince me that innovation needs to be steered away from automation.
- Strategies for an Accelerating FutureFour questions to ask your organization.
- Why Are LLMs So Gullible?It’s Because They’re Naive and Constantly Confused
- Google's Gemini Advanced: Tasting Notes and ImplicationsAnd then there were two.
January 2024
- What Can be Done in 59 Seconds: An Opportunity (and a Crisis)Five analytical tasks in under a minute
- Anatomy of a HackPractical hacks are complicated. That means complicated systems are unsafe.
- Will AI transform law?The hype is not supported by current evidence
- Generative AI’s end-run around copyright won’t be resolved by the courtsOutput similarity is a distraction
- What is “Prompt Injection”, And Why Does It Fool Chatbots?LLMs See Words, But Not Context
- The Lazy Tyranny of the Wait CalculationTaking AI timelines seriously
- Signs and PortentsSome hints about what the next year of AI looks like
2023
December 2023
- Will scaling work?Data bottlenecks, generalization benchmarks, primate evolution, intelligence as compression, world modelers, and other considerations
- An AI Haunted WorldIntelligence, everywhere.
- Would We Really Shut Down A Misbehaving AI?Ask The Old OpenAI Board How Things Went When They Tried To Shut Down Sam Altman
- Are open foundation models actually more risky than closed ones?A policy brief on open foundation models
- An Opinionated Guide to Which AI to Use: ChatGPT Anniversary EditionA simple answer, and then a less simple one.
- Model alignment protects against accidental harms, not intentional onesThe hand wringing about failures of model alignment is misguided
November 2023
- Reshaping the tree: rebuilding organizations for AITechnological change brings organizational change.
- Toward Better AI MilestonesHow to get off the treadmill of constantly shifting goalposts, and define tests that will tell us when AI capabilities get serious
- Not much is changing, a lot is changingOpenAI, Microsoft, and the OpenOffspring
- Almost an Agent: What GPTs can doAlso, my book has a cover (also I have a book coming out)
- Working with AI: Two paths to promptingDon't overcomplicate things
October 2023
- What the executive order means for openness in AIGood news on paper, but the devil is in the details
- The Best Available Human StandardWhat are the imperatives of the upside?
- How Transparent Are Foundation Model Developers?Introducing the Foundation Model Transparency Index
- What people ask me most. Also, some answers.A FAQ of sorts
- Scale, schlep, and systemsThis startlingly fast progress in LLMs was driven both by scaling up LLMs and doing schlep to make usable systems out of them. We think scale and schlep will both improve rapidly.
- Evaluating LLMs is a minefieldAnnotated slides from a recent talk
- The shape of the shadow of The ThingWe can start to see, dimly, what the near future of AI looks like.
September 2023
- When Will AIs Acquire Insight?How Could We Even Train Them For That?
- Everyone is above averageIs AI a Leveler, King Maker, or Escalator?
- The AI Explosion Might Never HappenAs AI Learns To Improve Itself, Improvements Will Become Harder To Find
- Centaurs and Cyborgs on the Jagged FrontierI think we have an answer on whether AIs will reshape work....
- Intermediate SuperintelligenceGodlike AI, if it happens at all, won't happen overnight
- Embracing weirdness: What it means to use AI as a (writing) toolAI is strange. We need to learn to use it.
- When We Forecast AGI, Do We Mean In The Lab Or In The Field?The Onset of AGI Will Play Out Over Decades
August 2023
- Language models surprised usMost experts were surprised by progress in language models in 2022 and 2023. There may be more surprises ahead, so experts should register their forecasts now about 2024 and 2025.
- Now is the time for grimoiresIt isn't data that will unlock AI, it is human expertise
- Does ChatGPT have a liberal bias?A new paper making this claim has many flaws. But the question merits research.
- The AI Progress ParadoxReceding Horizons and Misleading Milestones
- Introducing the REFORMS checklist for ML-based scienceML-based science is in trouble. Clear reporting standards for researchers could help.
- Automating creativityThere is now strong evidence that AI can help make us more innovative.
- ML is useful for many things, but not for predicting scientific replicabilityHow the veneer of AI is used to legitimize awful ideas
- Why We Won't Achieve AGI Until Memory Is A Core Architectural ComponentThe Hard Part of Thinking Is Knowing What To Think About
- In Praise of Boring AIAutomation has always been about killing tedious work. AI can do the same.
July 2023
- On holding back the strange AI tideThere is no way to stop the disruption. We need to channel it instead
- Is GPT-4 getting worse over time?A new paper going viral has been widely misinterpreted
- Process vs. Product: Why We Are Not Yet On The Cusp Of AGIYou Can't Learn The Journey By Observing The Destination
- How to Use AI to Do Stuff: An Opinionated GuideCovering the state of play as of Summer, 2023
- We Need to Recognize How Profoundly Different The AGI Future Will BeWe are shaping AI; soon it will also be shaping us
- What AI can do with a toolbox... Getting started with Code Interpreter [Now called Advanced Data Analytics]Democratizing data analysis with AI
- The Homework ApocalypseFall is going to be very different this year. Educators need to be ready.
June 2023
- Generative AI companies must publish transparency reportsThe debate about the harms of AI is happening in a data vacuum
- On giving AI eyes and earsAI can listen and see, with bigger implications than we might realize.
- Three Ideas for Regulating Generative AIPolicy input to the federal government from a Stanford-Princeton team
- Is AI-generated disinformation a threat to democracy?An essay on the future of generative AI on social media
- Detecting the Secret CyborgsThe AI Trap for Organizations
- What Will AI Do For Us In The Near Term?The more time I spent writing this post, the more optimistic I became
- Contra Marc Andreessen on AI"The claim that you will completely control any system you build is obviously false, and a hacker like Marc should know that"
- Assigning AI: Seven Ways of Using AI in ClassAlso prompts! And things to watch out for!
- Why trying to "shape" AI innovation to protect workers is a bad ideaInstead, we should empower workers and create mechanisms for redistribution.
- Licensing is neither feasible nor effective for addressing AI risksNon-proliferation only benefits incumbents
- Could AI accelerate economic growth?Most new technologies don’t accelerate the pace of economic growth. But advanced AI might do this by massively increasing the research effort going into developing new technologies.
- Setting time on fire and the temptation of The ButtonWe used to consider writing an indication of time and effort spent on a task. That isn't true anymore.
May 2023
- Is Avoiding Extinction from AI Really an Urgent Priority?The history of technology suggests that the greatest risks come not from the tech, but from the people who control it
- To Address AI Risks, Draw Lessons From Climate ChangeWe can't solve the problem today; we *can* start establishing the conditions for it to be solved.
- What happens when AI reads a book 🤖📖And some prompts that might be useful when it does.
- A Unified Theory of AI RiskNear-term risks aren't a distraction from extinction scenarios, they're the practice exam
- On-boarding your AI InternThere's a somewhat weird alien who wants to work for free for you. You should probably get started.
- Catastrophe / EucatastropheWe have more agency over the future of AI than we think.
- AI is not good software. It is pretty good people.A pragmatic approach to thinking about AI
- Get Ready For AI To Outdo Us At EverythingWe have time to prepare; step one is acknowledging what we're preparing for
- It is starting to get strange.Let's talk about ChatGPT with Code Interpreter & Microsoft Copilot
- The costs of cautionIf you thought we might be able to cure cancer in 2200, then I think you ought to expect there’s a good chance we can do it within years of the advent of AI systems that can do the research work humans can do.
April 2023
- Beyond the Turing TestThe key question about AI is no longer "can it think", but rather "can it hold down a job"?
- A guide to prompting AI (for what it is worth)A little bit of magic, but mostly just practice
- Quantifying ChatGPT’s gender biasBenchmarks allow us to dig deeper into what causes biases and what can be done about it
- An Intuitive Explanation of Large Language ModelsInstead of diving into the math, I explain *why* they're built as "predict the next word" engines, and present a theory for why they make conceptual errors.
- Democratizing the future of educationWe are all EdTech designers, now
- I’m a Senior Software Engineer. What Will It Take For An AI To Do My Job?Exploring the gaps between current LLMs and true general intelligence
- One sentence.Prompting for maximum impact (and why that is a bad idea)
- I set up a ChatGPT voice interface for my 3-year old. Here’s how it went.Chatbots are likely to revive familiar debates about kids and apps
- GPT-4 Doesn't Figure Things Out, It Already Knows ThemYes, It's a Stochastic Parrot, But Most Of The Time You Are Too, And It's Memorized a Lot More More Than You Have
- Continuous doesn’t mean slowOnce a lab trains AI that can fully replace its human employees, it will be able to multiply its workforce 100,000x. If these AIs do AI research, they could develop vastly superhuman systems in under a year.
- Nobody knows how many jobs will "be automated"Whatever that even means.
- The future of education in a world of AIA positive vision for the transformation to come
- Thinking companion, companion for thinkingSome simple ways to use AI to break you out of biases
- AIs accelerating AI researchResearchers could potentially design the next generation of ML models more quickly by delegating some work to existing models, creating a feedback loop of ever-accelerating progress.
March 2023
- Is it time for a pause?The single most important thing we can do is to pause when the next model we train would be powerful enough to obsolete humans entirely. If it were up to me, I would slow down AI development starting now — and then later slow down even more.
- A misleading open letter about sci-fi AI dangers ignores the real risksMisinformation, labor impact, and safety are all risks. But not in the way the letter implies.
- How to use AI to do practical stuff: A new guidePeople often ask me how to use AI. Here's an overview with lots of links.
- "Aligned" shouldn't be a synonym for "good"Perfect alignment just means that AI systems won’t want to deliberately disregard their designers' intent; it's not enough to ensure AI is good for the world.
- Alignment researchers disagree a lotMany fellow alignment researchers may be operating under radically different assumptions from you.
- The ethics of AI red-teamingIf we’ve decided we’re collectively fine with unleashing millions of spam bots, then the least we can do is actually study what they can – and can’t – do.
- Situational awarenessAI systems that have a precise understanding of how they’ll be evaluated and what behavior we want them to display will earn more reward than AI systems that don’t.
- Playing the training gameWe're creating incentives for AI systems to make their behavior look as desirable as possible, while intentionally disregarding human intent when that conflicts with maximizing reward.
- Training AIs to help us align AIsIf we can accurately recognize good performance on alignment, we could elicit lots of useful alignment work from our models, even if they're playing the training game.
- Superhuman: What can AI do in 30 minutes?AI multiplies your efforts. I found out by how much...
- Acceleration.7 days of new AI technologies shows us that everything is happening very fast.
- OpenAI’s policies hinder reproducible research on language modelsLLMs have become privately-controlled research infrastructure
- GPT-4 and professional benchmarks: the wrong answer to the wrong questionOpenAI may have tested on the training data. Besides, human benchmarks are meaningless for bots.
- Using AI to make teaching easier & more impactfulHere are five strategies and prompts that work for GPT-3.5 & GPT-4
- What is algorithmic amplification and why should we care?A symposium and a primer on social media recommendation algorithms
- How to... use AI to unstick yourselfWe often lose momentum because of something small. AI can help.
- Artists can now opt out of generative AI. It’s not enough.Opting out is the latest example of generative AI developers externalizing costs.
- LLMs are not going to destroy the human raceIt's just a chatbot, dude.
- Secret Cyborgs: The Present Disruption in Three PapersThe future is already here, we just need to figure out a few details.
- The LLaMA is out of the bag. Should we expect a tidal wave of disinformation?The bottleneck isn't the cost of producing disinfo, which is already very low.
- Feats to astonish and amazeA compendium of things I didn't think AI should be able to do
- Power and Weirdness: How to Use Bing AIBing AI is a huge leap over ChatGPT, but you have to learn its quirks
- AI cannot predict the future. But companies keep trying (and failing).A new paper on how AI companies make false promises and how we can challenge them
February 2023
- AI Techies!Everything you didn't realize you wanted to know about the denizens of Cerebral Valley
- How to Get an AI to Lie to You in Three Simple StepsI keep getting fooled by AI, and it seems like others are, too.
- Blinded by AnalogiesWhat is this AI thing? The wrong model can lead us astray
- People keep anthropomorphizing AI. Here’s whyCompanies and journalists both contribute to the confusion
- The future, soon: what I learned from Bing's AIWe had a brief glimpse of two different types of AI. Both are significant
- My class required AI. Here's what I've learned so far.(Spoiler alert: it has been very successful, but there are some lessons to be learned)
- I hope you weren't getting too comfortable.I just got access to the new Bing AI. My initial thoughts are that our assumptions about the limits of AI were wrong.
- A quick and sobering guide to cloning yourselfIt took me a few minutes to create a fake me giving a fake lecture.
- Magic for English MajorsProgramming in prose in an AI-haunted world
- "Do not fear AI, puny humans... that is not meant as a threat."What we can learn from a completely AI written & illustrated lecture
- The Machines of Mastery"Anyone can learn anything they want..." and how technology can help
January 2023
- A prosthesis for imagination: Using AI to boost your creativityAI can already beat humans in many measures of creativity. Let's use that to our advantage.
- The practical guide to using AI to do stuffA resource for students in my classes (and other interested people).
- Becoming strange in the Long SingularityA sudden increase in AI capability suggests a weirder world in our near future
- All my classes suddenly became AI classesWe can't beat AI, but it doesn't need to beat us (or our students)
- Secretary jobs in the age of AIA guest post by Hollis Robbins
- How to... use ChatGPT to boost your writingThe key to using generative AI successfully is prompt-crafting
- And the great gears begin to turn again...Progress in two vital areas had slowed, AI could change that
- You can't brute force the unsolvableBut a little finesse can help
- The third magicA meditation on history, science, and AI
2022
December 2022
- Who believes more myths about humans: AI or educated humans?ChatGPT should do badly where human information is most wrong. But is that true?
- The street finds its own uses for things, AI EditionBreakthrough innovations often come from people who use technology, not create it. So, how are students using AI?
- ChatGPT is my co-founderCommon barriers hold entrepreneurs back, AI can help founders cross them
- How to... use AI to teach some of the hardest skillsWhen errors, inaccuracies, and inconsistencies are actually very useful
- Four Paths to the RevelationGive me 10 minutes, and I think I can make you obsess over AI
- ChatGPT is a bullshit generator. But it can still be amazingly usefulThe philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever. Large Language Models (LLMs) are trained to produce
- The Mechanical ProfessorI take a job I know well, and try to see how far I can automate it with AI.
- How to... use AI to generate ideasAI can make you more creative right now, and here are some ways to do it...
- Generative AI: autocomplete for everythingA joint blog post by Noah and roon on the future of work in the age of AI
November 2022
- AI has a strategy.And we can learn a lot from how it is going to win.
- The bait and switch behind AI risk prediction toolsToronto recently used an AI tool to predict when a public beach will be safe. It went horribly awry. The developer claimed the tool achieved over 90% accuracy in predicting when beaches would be safe to swim in. But the tool did much worse: on a majority of the days when the water was in fact…
October 2022
September 2022
- Eighteen pitfalls to beware of in AI journalismA checklist for avoiding hype
- Lovecraftian intelligenceIs AI an incomprehensible cosmic horror?
- Generative AI models generate AI hypeOver 90% of images of AI produced by a popular image generation tool contain humanoids
- American workers need lots and lots of robotsWith the power of automation, our workers can win. Without it, they're in trouble.