Quick Facts

  • OpenAI cut GPT-5.6 Luna pricing by 80%, dropping the combined input/output cost from $7 to $1.40 per million tokens.
  • The cuts came just three weeks after the GPT-5.6 family reached general availability on July 9, 2026.
  • Chinese-origin models grew from under 2% of token consumption in late 2024 to more than 50% by June 2026, according to OpenRouter data.

OpenAI cut the price of its GPT-5.6 Luna model by 80% on July 30, slashing the combined input and output cost from $7 per million tokens to $1.40. The company also reduced pricing on its mid-tier GPT-5.6 Terra model by roughly 20%, while leaving its flagship Sol model unchanged.

CEO Sam Altman announced the changes on X, calling them “major price cuts.” The repricing landed just three weeks after the GPT-5.6 family reached general availability, an unusually short window that reflects the speed of competition now bearing down on AI model providers.

Where Luna Now Sits

At $1.40 per million combined tokens, Luna now undercuts Google’s Gemini 3.5 Flash-Lite, which costs $2.80, and sits far below Gemini 3.6 Flash at $9. Terra, now priced at $2 per million input tokens and $12 per million output tokens, comes in below Anthropic’s Claude Sonnet 4.6, which costs $3 input and $15 output.

OpenAI says Luna is built for high-volume tasks: large-scale document analysis, customer interaction sorting, and routine coding inside agent loops. The company claims the model matches frontier-class systems from a year ago at roughly 6 cents on the dollar per task and at nearly nine times the speed, based on its own testing. Third-party firm Artificial Analysis says Luna outperforms Gemini 3.6 Flash and the older Gemini 3.1 Pro.

Enterprise Customers Are Pushing Back on Costs

Altman acknowledged the pressure directly at a customer event in June. “Every enterprise now is thinking about spend and the value they’re getting in exchange for AI,” he said. He called costs “a huge issue” that had gone from never coming up to dominating client conversations within months.

The numbers back that up. Uber consumed its full 2026 AI coding budget in four months. Per-engineer costs for tools like Claude Code and Cursor ranged from $500 to $2,000 a month, prompting Uber to cap usage at $1,500 monthly per tool. Enterprises are shifting away from flat subscriptions toward usage-based pricing, which creates unpredictable bills as adoption scales.

Notion AI product lead Hoda Noorian said Terra matched GPT-5.5 quality on scoped tasks at half the cost and in 60% less time. Dust co-founder Stanislas Polu said Luna handles the same agentic work 40% faster and 40% cheaper than his firm’s previous default model.

The Competitive Math

OpenAI’s decision to gut Luna’s price while keeping Sol unchanged points to a deliberate two-track approach. The company is trying to hold its position as the provider of the most capable frontier models while making its utility tier cheap enough to stop enterprise volume from moving to lower-cost alternatives.

The threat from Chinese providers is concrete. OpenRouter data shows Chinese-origin models grew from under 2% of token consumption in late 2024 to more than 50% by June 2026. Moonshot AI’s Kimi K3, a 2.8-trillion-parameter model released earlier this month, has narrowed the performance gap with leading U.S. systems while giving developers the option to run models on their own infrastructure.

The competitive pressure also hits Anthropic hard. Anthropic’s annualized run rate jumped from $9 billion at the end of 2025 to $47 billion by May 2026, driven largely by Claude Code. The company posted its first profitable quarter in Q2 2026. But its models sit at the costlier end of the market, and OpenAI’s repricing narrows the gap further.

The Financial Strain Is Real

OpenAI posted a -122% adjusted operating margin in Q1 2026, meaning it lost $1.22 for every dollar it brought in. Analysts warn that cheaper tokens could lift usage volumes while deepening financial strain, especially as both OpenAI and Anthropic filed confidential listing prospectuses in June.

OpenAI also added a Fast mode for Sol through the API, delivering up to 2.5 times standard processing speed at double the price. The feature replaces the company’s previous Priority Processing offering.

Read more: AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.