Quick Facts

  • GPT-6 Sol is priced at $2 per million input tokens and $10 per million output tokens, down 50% from GPT-5.6 Sol pricing.
  • GPT-6 Luna's output price drops 58%, from $1.20 to $0.50 per million output tokens, with input tokens now at $0.10 per million.
  • OpenAI overhauled its prompt caching system, offering a 90% discount on cached input-token reads and cutting fresh token processing by more than 50% across billions of requests.

OpenAI released two new models on Sept. 22, 2026: GPT-6 Sol and GPT-6 Luna. Both sit below the flagship GPT-6 Astra in capability and price, and both carry permanent price cuts of at least 50% compared with their GPT-5.6 predecessors.

The company said the cuts are not promotional. OpenAI attributed the savings to improvements in caching and inference efficiency, stating it is "passing those savings directly on to users and customers."

Where the Models Fit

GPT-6 Astra, released Sept. 3, anchors the high end at $10 per million input tokens and $50 per million output tokens, with a 1,050,000-token context window. Sol and Luna fill out the mid-tier and budget tiers below it.

Sol targets enterprise workflows that need strong accuracy at moderate cost. Luna is built for high-volume, lower-complexity tasks where price per call matters most.

Benchmark Results

On OpenAI's internal factuality evaluation, GPT-6 Sol makes roughly half as many errors as its predecessor and approaches Astra-level reliability. Luna, at higher effort levels, matches GPT-5.6 Sol at about one-hundredth the cost.

On AutomationBench 1.0.6, Sol at its highest effort setting scored 33.2% at $0.27 per task. Anthropic's Claude Opus 5 at max effort scored 26.9% at 11.1 times that cost. Luna scored 5.4 percentage points higher than GPT-5.6 on the same benchmark while costing 58% less per task.

For computer use tasks on OSWorld 2.0, Sol nearly matched Claude Opus 5's medium-effort score, 60.5% to 60.3%, at roughly 80% lower cost. On Agents' Last Exam, which tests AI agents across 55 sub-industries, Sol at highest effort scored 56.4%, which OpenAI says beats Claude Opus 5's best score at 60% lower cost per task.

Software engineer Hitesh Rohira noted that on DeepSWE v1.1, GPT-6 Luna at maximum reasoning effort reaches 66.6%, edging past GPT-6 Sol at high effort while costing a fraction of the price.

Caching Overhaul

Alongside the model releases, OpenAI announced a major update to its caching infrastructure. Cached input-token reads now carry a 90% discount, up from previous rates. New tools include a Prompt Caching Dashboard, a diagnostics tool that flags missed caching opportunities, and explicit breakpoints that let developers control where cached prompt prefixes end.

Developers can now change reasoning effort or toggle tools without losing cached context. GitHub said the improvements reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models over recent months, with measurable speed gains for Copilot users.

Competitive Timing

Anthropic released a new version of Claude Opus 5.5 roughly 90 minutes before OpenAI's announcement on Sept. 22. The back-to-back releases reflect the pace of competition between the two companies heading into OpenAI's DevDay event on Sept. 29.

CEO Sam Altman has said the company's goal is to "offer leading AI models across every price point and modality." The permanent pricing structure for Sol and Luna gives developers a stable cost basis for production applications, removing the uncertainty that comes with time-limited discounts.

Read more: OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.