Threads

Generative media

The rise of AI image, video and audio generation — from DALL·E and Stable Diffusion to Sora and Veo 3 — and the misuse, from nonconsensual deepfakes to copyright disputes, that tracked its quality.

Generative media is the thread of models that make images, video and audio rather than text. It opened with OpenAI’s DALL·E and CLIP, the pairing that connected language to pictures, and accelerated through DALL·E 2, Midjourney’s Discord beta and the open release of Stable Diffusion, which put photorealistic image generation in anyone’s hands.

Video was the next frontier. OpenAI previewed Sora in early 2024 and released it publicly that December; Google answered with Veo 3, and OpenAI’s native image generation in GPT-4o set off a viral Studio Ghibli-style wave. By late 2025 Sora 2 had turned generation into a social feed.

Quality brought misuse in step. The same tools were used to make nonconsensual sexual deepfakes, and xAI’s image features were turned to mass-produced abuse imagery — harms that put generative media at the centre of the deepfake-law and consent debates traced in other threads. The tension is constant across it: each gain in fidelity widened both the creative use and the capacity for harm.