Opus 5.5 vs. GPT-6 Sol and Luna: the frontier price war arrives
Anthropic and OpenAI shipped new flagship models within an hour of each other on September 22, and both led with price. Here's what changed and how we'd route work across them.

Two frontier labs shipped on the same day, and neither led with a benchmark. On September 22, Anthropic released Claude Opus 5.5. Roughly an hour later OpenAI answered with GPT-6 Sol and GPT-6 Luna. The headline for both was cost.
What shipped
Claude Opus 5.5 is Anthropic's new flagship. List pricing is $4 per million input tokens and $20 per million output tokens, about 20% below the $5 / $25 of earlier Opus models. Cache reads dropped 60%, the biggest single cut on the price sheet. Anthropic also says the model responds about 30% faster, with a 1M-token context window, up to 128K output tokens, and adaptive thinking that's always on.
GPT-6 Sol is OpenAI's coding and agent model. GPT-6 Luna is the budget option for high-volume work like summarisation and extraction. Both accept text and images and return text, and both have the same unusually large limits: a 1,050,000-token context window and up to 128,000 output tokens. Simon Willison notes the GPT-6 prices are roughly half those of the GPT-5.6 equivalents.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 |
| Grok 4.7 | $2.00 | $0.50 | $6.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
The fine print: more thinking isn't always better
Early hands-on testing turned up one caveat worth knowing. Willison found that Opus 5.5 at its maximum reasoning level could run past its 128,000-token output limit while working on a simple drawing task, and return nothing. Fable 5.1's maximum mode didn't have the problem.
A model that thinks until it runs out of room gives you a bill and no answer.
For production agents, that's a reminder to set reasoning effort per task, not globally, and to treat a truncated response as its own failure case with its own retry path.
How we'd route work across them
Cheaper frontier tokens don't flatten the market. They make routing matter more. Here's roughly where we'd start a client in October:
- Luna for the long tail. Classification, extraction, triage and summarisation at $0.50 per million output tokens change the unit economics of "run an agent on every ticket."
- Sol or Opus for the hard steps. Planning, multi-file code changes and tool-heavy workflows still reward the bigger models. At these prices, cutting capability to save money is a much weaker argument than it was a quarter ago.
- Lean on the cache. Both vendors cut cached-input pricing hard. Agents that reuse long system prompts, tool schemas and retrieved context should be built to hit the cache on purpose.
- Re-run your evals, not the leaderboards. A same-day launch from two vendors is the perfect moment to run your own task suite against both. The model that wins on your data is the one to ship.
The bottom line
The frontier is getting cheaper faster than it's getting smarter. For teams already in production, the quickest win this quarter probably isn't a new capability. It's re-pricing the agents you already run.


