DeepMind controls a fusion plasma with reinforcement learning
A single network commanding all of a tokamak's control coils held plasma shapes on Switzerland's TCV reactor, including configurations conventional controllers struggle with.
- Models & capabilities
- Ideas & essays
- Notable
DeepMind, working with the Swiss Plasma Center at EPFL, published a Nature paper describing a reinforcement-learning system that controlled the magnetic coils confining plasma inside a working nuclear fusion reactor, the Variable Configuration Tokamak (TCV) in Lausanne. Rather than the bespoke feedback controllers tokamaks normally rely on for each configuration, a single neural-network policy learned to command the reactor’s full set of control coils directly, meeting shape and position targets specified at a high level while respecting the device’s physical and operational constraints.
The system was trained inside a simulator of the tokamak and then deployed directly onto the physical hardware without further tuning — DeepMind described this as successfully bridging the “sim-to-real” gap, a step that has proven difficult in most other robotics and control applications of reinforcement learning. On the real reactor, the controller produced and held a range of plasma shapes, including standard elongated configurations as well as harder ones, such as “negative triangularity” and “snowflake” geometries that specialised controllers typically struggle with. It also sustained two separate plasma “droplets” simultaneously within the vessel, a configuration not previously demonstrated this way.
The result did not represent progress in achieving fusion ignition or net energy gain — TCV is a research tokamak used to study plasma physics and control, not a power-generating device — and the paper’s contribution was narrowly about control, not about the reactor’s core physics. DeepMind and EPFL framed the work as a tool for physicists: a controller flexible enough to let researchers explore plasma configurations quickly rather than engineering a dedicated feedback loop for every shape.
It was among the earliest widely covered examples of a national fusion research programme adopting a general-purpose learned controller in place of classical control engineering, and previewed a wave of subsequent AI-for-science work applying similar techniques to reactor control and plasma diagnostics at other facilities.