Meta unveils the AI Research SuperCluster
A 6,080-GPU first phase, to reach 16,000 GPUs and a claimed 5 exaflops of mixed-precision compute by mid-2022, built for training on trillions of examples.
- Compute & infrastructure
- Labs & people
- Notable
Meta announced the AI Research SuperCluster (RSC), a computing cluster it described as one of the fastest AI supercomputers running anywhere and, once complete, the fastest in the world. The first phase comprised 760 NVIDIA DGX A100 systems — 6,080 A100 GPUs — linked over an NVIDIA Quantum InfiniBand fabric with no oversubscription, backed by Pure Storage flash arrays totalling more than 230 petabytes.
Meta said the cluster would grow to 16,000 GPUs by mid-2022, at which point it projected close to 5 exaflops of mixed-precision compute. The company gave comparative figures for early workloads: computer-vision pipelines running roughly 20 times faster than on its previous infrastructure, and large NLP models — tens of billions of parameters — training in about a third of the time. RSC replaced infrastructure Meta had built during the pandemic and largely operated remotely.
The stated purpose was training progressively larger models on progressively larger datasets, with Meta saying it aimed eventually to train models on trillions of examples across dozens of languages, using not just labelled data but also unlabelled and multimodal examples “sourced from Meta’s production systems.” The announcement tied this directly to the company’s metaverse ambitions, framing RSC as infrastructure for real-time voice translation across large multilingual groups and other AI systems Meta said it needed to build the next computing platform.
RSC arrived a year before ChatGPT reset public attention on language models, and before “GPU cluster size” became a routine unit of competitive comparison between labs. It is one of the earliest entries in what became a broader pattern: computing capacity announced in advance of, and partly to justify, the scale of the training runs a lab intended to run — a pattern OpenAI, Microsoft, Google and xAI would each repeat at larger scale over the following years.