GPT-6 Sol, Luna, or Claude Opus 5.5? Compare the Work You Can Actually Use

OpenAI and Anthropic released three models on September 22, 2026. The price tags are easy to quote; the cost of a report your team can actually send is harder to see. If an inexpensive draft needs two reruns and an hour of fact-checking, the cheapest token is not necessarily the cheapest finished task.
What changed on September 22
OpenAI introduced GPT-6 Sol for demanding coding and agentic work and GPT-6 Luna for focused, high-volume tasks. Its launch post says API prices are 50% below GPT-5.6 promotional prices. At publication, OpenAI lists Sol at $2 input and $10 output per million tokens, and Luna at $0.10 and $0.50. Anthropic introduced Claude Opus 5.5 on the same day. It lists $4 input and $20 output per million tokens, a 20% reduction from Opus 5; its $0.20 cache-read price is 60% lower than before. Anthropic's estimate of 40% lower cost on typical workloads is a vendor estimate, not another reduction in token list price. Caching, reasoning settings and tools can change the bill.
Why the token price is only the first line of the budget
Consider an illustrative job that uses 100,000 input tokens and produces 10,000 output tokens, with no cache or tools. At the standard listed rates checked September 24, the model charge is about $0.30 for GPT-6 Sol, $0.015 for GPT-6 Luna, and $0.60 for Claude Opus 5.5. This is arithmetic on a hypothetical workload, not a measured result or a fair quality comparison. If a Luna draft needs repeated runs, or an Opus draft saves an hour of review, the answer changes. Even a sound first-pass estimate can be distorted by the mix of input and output: output tokens cost more than input tokens for all three models.
OpenAI's Sol and Luna documentation also says requests above 272,000 input tokens are billed at higher rates for the entire request, despite the much larger advertised context window. Fast processing and regional processing have separate price adjustments. Anthropic highlights cheaper cache reads for Opus 5.5; the benefit depends on whether a team repeatedly reuses the same context. Those are procurement details for an actual workflow, not reasons to choose a model solely from a headline price.
What the benchmarks actually measure
Artificial Analysis reported Opus 5.5 at 58 on its Intelligence Index at max effort, its highest measured score at the time. That is useful independent evidence about a specified evaluation. It is not a head-to-head verdict on your team's supplier memo, sales analysis or customer brief. Sol and Luna have a different price/performance proposition, and a result obtained at one reasoning setting should not be placed next to another setting as though they were the same test. A long context window also does not certify that every figure in a long document was found and cited correctly.
Suppose your team prepares a weekly competitor update from three product pages, two PDFs and a pricing spreadsheet. The deliverable is a one-page recommendation with links, a change log and numbers that reconcile. A model can produce fluent prose while missing the changed price in a footnote. That miss matters more than how impressive the first paragraph sounds.
A realistic split across three tasks
For a weekly report that synthesizes conflicting evidence, track whether the model preserves source links and flags unresolved numbers. For a high-volume extraction job, count field-level errors and reruns; low per-call cost can be compelling here. For a presentation outline, check whether the recommendations follow from the supplied data and whether an editor can use the structure without rewriting it. The same model need not win all three. A decision memo can recommend one default, one exception path and a date to retest when the workload or prices change.
The most informative table is therefore not just “price / score.” Use columns for input packet, accepted-result criteria, first-pass success, rerun count, review minutes, provider charges and remaining risks. That table makes a cheaper model's strengths visible without hiding its correction costs—and gives a higher-priced model a chance to justify its place through a better finished result.
Measure cost per accepted deliverable
Run each candidate on the same three to five anonymized tasks, with the same input files and an agreed output format. Record whether the draft passes the actual acceptance check: correct figures, traceable sources, required sections, no unsupported assertions and edits a colleague can make. Then include API calls, tool charges, failed attempts, revisions and human review time in the cost of each accepted result. If one model needs a second pass on half the tasks, calculate that cost rather than treating every first answer as a success.
A useful decision memo separates three outcomes: use a stronger model for exception-heavy work; use a lower-cost model for repeatable extraction; or keep the current setup until the gap is meaningful. It should include examples of failures, not just an average score.
Turn the comparison into a decision brief
Give MuseWork the public announcements, your approved sample materials and the criteria your team uses to sign off a deliverable. Ask it to organize the claims, identify mismatched testing conditions and draft a small trial plan and reviewable decision brief. A ready-to-use request: “Using these model announcements, invoices and three anonymized briefs, create a comparison matrix for source accuracy, accepted-output rate, revision time and total cost. Flag what still needs a real test.”
For background on the earlier flagship release, read the GPT-6 Astra article.
From these numbers to your next deliverable
Bring your own notes, reports or source links to MuseWork and ask for a brief your team can review. Start a research task, or see how Deep Research works.
Keep the same work at hand when you leave your desk. Get MuseWork for iPhone or Android, or install MuseWork Desktop to work with local files with your permission.
Sources
- OpenAI model documentation: GPT-6 Sol and GPT-6 Luna, checked September 24, 2026.
- OpenAI launch announcement, September 22, 2026: Introducing GPT-6 Sol and Luna
- OpenAI API price list, checked September 24, 2026: API pricing
- Anthropic launch announcement, September 22, 2026: Introducing Claude Opus 5.5
- Artificial Analysis independent assessment, September 22, 2026: Independent assessment of Claude Opus 5.5

