On July 30, 2026, OpenAI announced steep price cuts for its lower-tier AI models, slashing GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. Enabled by GPU kernel and speculative decoding optimizations, these reductions aim to counter rising business cost concerns and intense competition from cheaper Chinese models like Moonshot AI's Kimi and Z.ai's GLM family.
GPT-5.6 Luna price cuts
- ▪The pricing for GPT-5.6 Luna decreased to $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6 respectively
- ▪OpenAI reduced the price of its GPT-5.6 Luna model by 80% on July 30, 2026
OpenAI inference optimization improvements
- ▪OpenAI stated that the price reductions were enabled by efficiency gains from GPT-5.6, including the model's ability to improve code and optimize performance during internal development
- ▪OpenAI achieved 20% lower serving costs from production GPU kernel improvements and 15% better token-generation efficiency from improved speculative decoding
Chinese AI competition pressure
- ▪US AI companies face pressure from Chinese rivals to improve inference efficiency and lower costs as businesses scrutinize ballooning AI spend
- ▪Chinese AI models, including Moonshot AI's Kimi and Z.ai's GLM family, offer strong capabilities at significantly lower prices than leading US models
GPT-5.6 Sol Fast mode
- ▪The new Fast mode for GPT-5.6 Sol delivers up to 2.5x faster speeds than Standard processing but costs twice as much
- ▪OpenAI introduced a new Fast mode for GPT-5.6 Sol in the API, replacing the Priority Processing option
ChatGPT Auto-review cost efficiency
- ▪OpenAI expects the Auto-review feature to cost approximately 10x less than before due to the transition to GPT-5.6 Luna
- ▪The Auto-review feature in the ChatGPT app and Codex CLI is now powered by GPT-5.6 Luna instead of GPT-5.4
Story comments
Loading comments…