Timeline

OpenAI releases GPT-4.5

Priced at $75/$150 per million tokens, about thirty times GPT-4o's rate, and retired from the API within five months in favour of the cheaper GPT-4.1.

  • Models & capabilities
  • Notable

OpenAI released GPT-4.5, which it described as its largest and most compute-intensive model to date, after a development process reported to have run behind schedule. Unlike o1, o3-mini and the reasoning-focused releases OpenAI had shipped since 2024, GPT-4.5 was a conventional pretrained model without a chain-of-thought reasoning stage — the last release in the GPT line built primarily by scaling unsupervised pretraining rather than the newer reasoning approach the company had begun prioritising.

OpenAI framed the release around broader world knowledge, better ability to follow user intent, and higher “EQ” — improved conversational warmth and creative writing — rather than reasoning-benchmark scores, where it lagged: by OpenAI’s own figures, GPT-4.5 trailed o3-mini, Claude 3.7 Sonnet and DeepSeek R1 on AIME and GPQA. Its clearest improvement was in factual reliability: on OpenAI’s SimpleQA test, the model gave a confabulated answer 37.1% of the time, against 59.8% for GPT-4o and 80.3% for o3-mini.

The model launched through the API at $75 per million input tokens and $150 per million output tokens — roughly thirty times GPT-4o’s rate — and OpenAI said the cost was significant enough that it was uncertain whether to keep serving the model long-term. Access began with ChatGPT Pro subscribers and paid API developers immediately, with Plus and Team tiers following within the week. A line in OpenAI’s accompanying documentation initially stated GPT-4.5 “is not a frontier model”; the phrase was removed after publication, and OpenAI stressed the release was a research preview rather than “a drop-in replacement for GPT-4o.”

The uncertainty proved warranted. OpenAI deprecated GPT-4.5 from its API in July 2025, less than five months after release, pointing developers to the cheaper GPT-4.1, and later withdrew it from ChatGPT entirely. It stands as the clearest instance of a frontier lab testing, and then abandoning, further scaling of pretraining alone once reasoning-focused training had become the more cost-effective route to capability gains.