AMD publishes a technical report for an open-weight model trained on its own GPUs
Instella-MoE, a 16-billion-parameter model with 2.8 billion active parameters, was pretrained entirely on AMD Instinct MI300X and MI325X chips rather than Nvidia hardware.
- Open weights & ecosystem
- Compute & infrastructure
- Minor
AMD published a technical report for Instella-MoE, a mixture-of-experts language model with 16 billion total parameters and 2.8 billion active per token, which the company said was trained entirely from scratch on its own Instinct MI300X and MI325X accelerators rather than Nvidia GPUs. AMD said it would release the full model flow — weights, training configuration, data mixtures and training code — to support reproducible research.
The significance is less the model’s capability, which the report positions as a mid-sized open-weight system rather than a frontier result, than the demonstration itself: that a complete pretraining run can be carried out end-to-end on AMD’s accelerators, which have historically trailed Nvidia’s in software maturity for large-scale training rather than inference. The release follows AMD’s acquisition of the inference-chip startup Taalas weeks earlier, part of a broader push to build a credible non-Nvidia stack across both training and serving.
For labs and governments concerned about a single-vendor dependency in AI compute, an open-weight model trained wholly on AMD hardware is a concrete data point rather than a roadmap promise, even though it does not by itself establish that AMD’s chips are competitive with Nvidia’s at frontier scale.