Meta pulls Galactica after three days
Its errors were formatted exactly like real citations and papers, so only an expert reader could tell fabrication from fact — a distinct failure mode from earlier chatbots' obvious mistakes.
- Models & capabilities
- Culture & impact
- Notable
Meta AI released Galactica, a large language model trained on scientific papers, reference material and knowledge bases, and put a public demo online alongside claims that it could write papers, literature reviews and Wikipedia-style articles complete with citations and formulas. Meta’s chief scientist Yann LeCun promoted it the same day, saying the demo could “generate a paper with relevant references, formulas, and everything.”
Within hours, researchers and other users found the demo readily produced confident, fluent, scientific-sounding text that was false — invented papers with plausible author names and citations that did not exist, fabricated data, and content criticised as biased. The problem was not that Galactica sometimes said something questionable; it was that its errors were formatted exactly like real scholarship, complete with the syntax of citation and method, so that only a reader already expert in the specific claim could tell fabrication from fact. Critics argued this made it more dangerous than a text generator with no scientific pretensions, since its output was designed to look authoritative.
Meta took the public demo offline after three days, saying in a statement that it had “paused the demo for now” in response to feedback. LeCun continued to defend the underlying research, tweeting sardonically three days after launch that the demo was “no longer possible to have some fun by casually misusing it.” The model’s weights and code remained available for research use under a Creative Commons licence.
The episode became a widely cited case study in the gap between a model performing well on benchmarks — Galactica’s paper reported strong scores on several scientific and mathematical evaluations — and a model being safe to expose to the public, arriving two weeks before ChatGPT’s launch made the same tension a mainstream concern rather than a research-community one.