OpenAI adds native image generation to GPT-4o
Images were generated natively by GPT-4o's own architecture rather than by a separate diffusion model, improving text rendering and editing of existing images.
- Models & capabilities
- Major
OpenAI built image generation directly into GPT-4o rather than routing requests to a separate model, as ChatGPT had done since 2023 with DALL·E 3. Because the same network that understood the prompt was the one producing the pixels, OpenAI said the system could follow detailed, multi-step instructions more reliably than before, render legible text inside images — long a weak point for diffusion-based generators — and edit or extend an uploaded photograph while preserving its content, rather than only generating from scratch.
The accompanying addendum to the GPT-4o system card described the safety measures attached to the launch, including C2PA metadata embedded in generated images to mark them as AI-produced and provide a verifiable record of origin, alongside internal tooling to help identify whether a given image came from OpenAI’s models. OpenAI said the feature would roll out gradually across ChatGPT’s Plus, Pro and free tiers, with wider API access to follow.
Demand outstripped what OpenAI had provisioned for launch. Within days the company imposed temporary rate limits — three generations a day for free-tier users — and delayed the wider free-tier rollout, with Sam Altman writing that “our GPUs are melting” and asking users to moderate how much they generated while the company worked on efficiency. The strain was driven substantially by a single trend: users converting photographs and film stills into the visual style of Studio Ghibli, a wave large enough to draw its own scrutiny and to become a separate flashpoint in the debate over AI and copyright.
The release marked a shift in how the leading chat products handled images — treating generation as a native capability of the core model rather than a bolted-on tool call — and rivals followed with their own natively integrated image systems over the following year.