OpenAI releases GPT-6 Sol and Luna and halves the API price of both tiers
GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, and GPT-6 Luna costs $0.10 and $0.50. OpenAI measures the 50% cut against the promotional prices it charged for GPT-5.6. The launch also brings a reworked prompt cache with discounts of up to 90% on cached input. Every benchmark comparison with Claude comes from OpenAI and uses Opus 5, not Opus 5.5.
OpenAI released two smaller members of the GPT-6 family on September 22, 2026. GPT-6 Sol replaces GPT-5.6 Sol, and GPT-6 Luna replaces GPT-5.6 Luna. In the API they are available as gpt-6-sol and gpt-6-luna. OpenAI says GPT-6 Astra, released on September 3, 2026, remains its strongest model.
What the two models cost
Sol is priced at $2 per million input tokens and $10 per million output tokens. Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens. The predecessors cost $4 and $20 for Sol and $0.20 and $1.20 for Luna, which OpenAI presents as a 50% cut on both tiers.
Two details sit behind that figure. OpenAI compares the new prices with what it calls GPT-5.6 promotional pricing, not with a standard price list. And Luna's output price fell by more than half: going from $1.20 to $0.50 is a 58% cut, by our calculation.
A cache that keeps shared context for 30 minutes
OpenAI also rebuilt prompt caching for the GPT-6 family. Cached input tokens are discounted by up to 90%, and a shared prompt prefix qualifies when it is reused within 30 minutes. The company says the new system hits the cache more often with no change on the developer's side.
The change matters most for agents, which resend the same instructions, tool definitions, and earlier turns with every request. On GPT-6 models, reasoning effort can now change between responses through a configuration_update message without losing the cache. OpenAI recommends keeping tool definitions stable and narrowing them with allowed_tools, or setting tool_choice to none, rather than deleting them.
GitHub, quoted by OpenAI, says that over several months the share of prompt tokens needing fresh processing fell by more than 50%, across billions of requests to OpenAI models.
- The Prompt Caching Dashboard shows what share of input is served from cache over time.
- A diagnostics tool compares a request with a recent response and names the change that caused a cache miss.
- Explicit breakpoints set where the cached part of a prompt ends.
- Prewarming processes known context, such as shared instructions, before the first user request arrives.
What OpenAI's own benchmarks claim
Every comparison in the announcement comes from OpenAI. On AutomationBench, which tests business workflows across 47 tools, the company reports 33.2% for Sol at xhigh effort, at $0.27 per task. By the same table, Claude Opus 5 at max effort scores 26.9% at 11.1 times the cost per task.
On DeepSWE v1.1, OpenAI puts Sol at 68.8% at max effort and Claude Fable 5 at 69.9% at xhigh effort, with Sol about 80% cheaper per task. Luna at max effort reaches 66.6%. On OSWorld 2.0, the company reports 60.5% for Sol at xhigh effort against 60.3% for Opus 5 at medium effort.
On an internal factuality test built from conversations in which users had flagged an error, OpenAI says Sol makes about half as many mistakes as its predecessor. The company notes that these conversations do not represent typical use.
What the comparison leaves out
Anthropic released Claude Opus 5.5 on September 22, 2026, at $4 per million input tokens and $20 per million output tokens. None of OpenAI's tables include it; Sol is compared with Opus 5, Fable 5.1, and Fable 5. At list prices per token, Sol costs half as much as Opus 5.5 on both input and output. This suggests the gap in cost per task is smaller than the tables show, since Opus 5.5 is itself cheaper than Opus 5, but no public measurement confirms that yet.
Where the models are available
In ChatGPT Work and Codex, both models are available to Plus, Pro, Business, Enterprise, and Edu users. Free and Go users get GPT-6 Luna in the desktop app. OpenAI says the models are not yet available in Chat and are rolling out gradually.
„Improvements in caching and inference let us serve these models at lower cost.“
Sources
Related

Anthropic releases Claude Sonnet 5.5, with input and output tokens at half the Opus 5.5 price
Claude Sonnet 5.5 keeps Sonnet 5's prices: $2 per million input tokens and $10 per million output tokens. That is half of what Opus 5.5 charges, while Anthropic's own benchmarks put the two models within a few points of each other. Code that turns thinking off needs a change before the switch, because the old setting now returns an error.
Anthropicverified
H Company releases Holo4, open-weight models for agents that operate a computer
H Company released two Holo4 models on September 28, 2026, with weights on Hugging Face. The larger one scores higher on the company's own tests but may not be used commercially. The smaller one is licensed under Apache 2.0.
Hugging Faceverified

Anthropic ships Claude Opus 5.5 and cuts token prices by a fifth
Opus 5.5 charges $4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Cache reads, which dominate the bill for agentic and coding work, fall 60% to $0.20 per million. Anthropic says the model works at the level of Claude Fable 5.1 on most tasks while costing 40% less to run than Opus 5. It also arrives with safeguards that reroute most cybersecurity work to Opus 4.8.
Anthropicverified

