Timeline

LAION releases Re-LAION-5B with CSAM links removed

The cleaned dataset shrank from 5.8 to 5.5 billion pairs after matching hashes from the Internet Watch Foundation and the Canadian Centre for Child Protection, eight months after Stanford's report.

  • Open weights & ecosystem
  • Security & misuse
  • Minor

LAION released Re-LAION-5B, a revised version of its LAION-5B image-text dataset, roughly eight months after Stanford Internet Observatory researchers reported finding links to suspected child sexual abuse material within it and LAION took the original dataset offline.

LAION said it removed 2,236 links to suspected CSAM in total, combining the 1,008 links Stanford had originally identified with 1,228 further matches found working with the Internet Watch Foundation and the Canadian Centre for Child Protection, which supplied hashes of known abuse material for comparison against the dataset’s image links. The organisation matched links by hash rather than by inspecting image content directly, and applied the same filtering to two dataset variants pegged to different safety thresholds. The cleaned dataset totalled 5.5 billion text-image pairs, down from 5.8 billion, and LAION described it as the first web-scale image-text dataset thoroughly checked against known CSAM hash lists.

The episode set a template other large web-scraped datasets faced pressure to follow: LAION-5B had underpinned major open image models including Stable Diffusion, so contamination anywhere in the pipeline raised the question of what downstream models trained on it had learned, not just what the dataset itself contained. The eight-month gap between the original report and the cleaned release reflected how slow and labour-intensive hash-based auditing was to carry out even with the cooperation of specialist child-safety organisations, at billion-scale.