This guide explains what GPT-6 Sol and Luna are, how they compare on reported benchmarks and pricing, and where they are available.
1. What Are GPT-6 Sol and GPT-6 Luna?
GPT-6 Sol and GPT-6 Luna are two models in OpenAI's GPT-6 family, designed to bring frontier intelligence to everyday work at different balances of capability and cost. They follow GPT-6 Astra, which OpenAI describes as its most intelligent and aligned model.
1.1 Why Did OpenAI Introduce Multiple GPT-6 Models?
OpenAI's reasoning is that work happens at different scales, rhythms, and budgets. The most demanding and important projects still call for Astra's full depth, but many tasks do not need that level of capability.
Sol and Luna were trained with methods similar to Astra's. The aim is to bring Astra's advances in professional work, factuality, coding, computer use, and alignment to faster, more affordable models. OpenAI describes this as advancing the frontier on cost efficiency, so advanced AI becomes practical for more everyday tasks and applications at scale.
2. Sol vs Luna vs Astra: How Do the Three Models Compare?
In short, Astra is OpenAI's highest-capability option, Sol is positioned for difficult work tasks with more room to iterate, and Luna carries the lowest listed prices in this announcement.
| Model | Positioning according to OpenAI | Price per 1M tokens (input / output) |
|---|---|---|
| GPT-6 Astra | Best model across the board, for the best results and an uncompromising experience | $10 / $50 |
| GPT-6 Sol | Takes on difficult work tasks, with higher usage limits and lower cost | $2 / $10 |
| GPT-6 Luna | Faster, more affordable model | $0.10 / $0.50 |
Astra's listed prices are five times Sol's on both input and output, and 100 times Luna's.
2.1 Which Types of Workloads Is Each Model Designed For?
OpenAI presents the models as tiers and does not assign fixed use cases to Sol and Luna, so this section sticks to what it stated.
- Astra: the most demanding and important projects, or when you want the best results and an uncompromising experience. OpenAI also calls it the world's best model for computer use.
- Sol: difficult work tasks, with more room to iterate through higher usage limits and lower cost.
- Luna: OpenAI describes Sol and Luna together as faster, more affordable models and does not name a separate use case for Luna. Its reported results show Luna scoring comparable to or above some other models on several tests at lower cost, such as DeepSWE and OSWorld 2.0.
Matching a model to a specific workload is an observation from these results, not an OpenAI recommendation, so teams should test on their own tasks.
3. Pricing and Token Economics
3.1 How Much Do GPT-6 Sol and Luna Cost?
Through the API, GPT-6 Sol costs $2 per 1 million input tokens and $10 per 1 million output tokens. GPT-6 Luna costs $0.10 per 1 million input tokens and $0.50 per 1 million output tokens.
Tokens are the small units of text a model reads (input) and writes (output), and API prices are quoted per 1 million of each.
OpenAI says improvements in caching and inference have reduced serving costs, and it is passing those savings on. GPT-6 Sol and Luna start at 50% lower input prices than their GPT-5.6 promotional counterparts, while Luna's output price is more than 58% lower.
| Model | Input | Output |
|---|---|---|
| GPT-5.6 Sol to GPT-6 Sol | $4 to $2 | $20 to $10 |
| GPT-5.6 Luna to GPT-6 Luna | $0.20 to $0.10 | $1.20 to $0.50 |
3.2 What Does That Look Like in Practice?
Consider a workload of 10 million input tokens and 2 million output tokens. On GPT-6 Sol it costs $40 ($20 for input plus $20 for output), compared with $80 at GPT-5.6 Sol prices. On GPT-6 Luna it costs $2.00, compared with $4.40 at GPT-5.6 Luna prices. This is an illustration of the listed prices, not a benchmark.
3.3 What Are Cost Per Task and Reasoning Effort?
Two terms appear throughout the results below. Reasoning effort is a setting that controls how much reasoning a model applies to a problem, and the labels low, medium, high, xhigh, and max refer to these levels. Cost per task is the cost of completing a single benchmark task, which lets readers compare models on both score and spend.
4. Professional Workflow Performance
4.1 How Do Sol and Luna Perform on Business Workflows?
On AutomationBench, AI agents are tested on end to end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. OpenAI reports that Sol at xhigh effort outperforms Claude Opus 5 at max effort at 9% of Opus 5's cost per task. Luna at high effort improves on its predecessor by 5.4 percentage points at 58% lower cost per task.
Reported AutomationBench results:
- Sol (xhigh): 33.2% at $0.27 per task
- Astra (low): 30.3%, at 3.9x Sol's cost
- Claude Opus 5 (max): 26.9%, at 11.1x
- Claude Fable 5.1 with Opus 5 fallback (max): 31.4%, at over 8.9x
OpenAI notes that the Fable 5.1 cost understates its actual cost, because it leaves out the cost of Opus 5 fallbacks, which occurred on about 40% of tasks.
4.2 What About Long Tasks Across Professional Fields?
Agents' Last Exam evaluates agents on long, economically valuable tasks spanning 55 sub-industries. Sol at max effort scores 56.4%, above Opus 5's highest score on the evaluation, at 60% lower cost per task.
5. Coding Performance
OpenAI reports that Sol improves substantially over GPT-5.6 Sol on FrontierCode and lands within about one point of Fable 5's top DeepSWE score at roughly 80% lower cost per task.
Coding agents now take on tasks with more complexity, scope, and duration. OpenAI says its internal daily token usage, valued at API prices, exceeds $600 for the median researcher and $7,000 at the 90th percentile. As tasks grow longer, the cost of sustained use matters more.
5.1 How Does FrontierCode Measure Coding Quality?
FrontierCode checks whether coding agents produce changes ready to merge into real codebases. Code is graded on correctness and on mergeability, which covers test quality, scope discipline, code style, and adherence to codebase standards. Sol matches Claude Fable 5.1 at xhigh effort at much lower cost.
5.2 What Do the DeepSWE Results Show?
DeepSWE v1.1 tests original, long software engineering tasks in real codebases.
- Sol at max effort scores 68.8%, within 1.1 points of Fable 5's highest score of 69.9% at xhigh effort, at approximately 80% lower cost per task.
- Luna at max effort scores 66.6%, comparable to Opus 5 and Fable 5 at medium effort. In these comparisons, Luna costs 93% less per task than Opus 5 and 96% less than Fable 5.
6. Computer-Use Capabilities
Astra remains OpenAI's best model for computer use, while Sol and Luna offer more cost efficient performance than their predecessors.
OSWorld 2.0 tests agents on long computer-use workflows spanning everyday and professional tasks. OpenAI reports partial reward on the offline set. Sol at xhigh effort scores 60.5%, against 60.3% for Opus 5 at medium effort, at approximately 80% lower cost per task. Luna at max effort exceeds GPT-5.6 Sol at medium effort at one tenth of the cost.
7. Factuality and Reliability
OpenAI's internal factuality evaluation uses de-identified real conversations in which users flagged mistakes by earlier OpenAI models. On it, Sol makes about half as many mistakes as its predecessor, approaching Astra level reliability at much lower cost. Luna also improves substantially, and at higher effort levels it matches GPT-5.6 Sol at about a hundredth of the cost.
OpenAI adds that these conversations are not representative of typical usage, where factual errors are rarer. Scores are not controlled for length, but its verbosity tests showed almost no dependence on answer length.
8. Collaboration and Response Style
Sol and Luna carry over Astra's improved communication style, which OpenAI expects to be most noticeable in technical and coding conversations. Expect more clarity, less jargon, fewer odd turns of phrase, fewer low value details, and slightly shorter answers without losing substance.
OpenAI illustrated this with a website design request. It said it preferred Sol's reply because it did not jump to conclusions, avoided restating obvious details, used less vague language, and was more open about what it had and had not checked. OpenAI also noted that style is subjective.
9. Prompt Caching for Agents and Long Conversations
Prompt caching lets a model reuse parts of a prompt it has already processed, which helps agents and long conversations that resend large amounts of context. OpenAI says GPT-6 has higher cache hit rates by default, with a 90% discount on cached input token reads.
Developers also get more ways to measure and improve caching:
- Monitor and diagnose: a Prompt Caching Dashboard shows how much input is cached over time, and a diagnostics tool explains missed caching opportunities.
- Adjust without breaking the cache: reasoning effort and tool availability can change mid conversation while preserving earlier context for reuse.
- Choose cached prefixes: explicit breakpoints let developers decide where cached prompt prefixes end.
GitHub reports that these improvements reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests, helping Copilot respond faster.
10. Alignment Improvements
Sol and Luna build on the alignment work introduced with Astra. In OpenAI's evaluations, both show improvements over their GPT-5.6 counterparts, including lower rates of misleading claims about their coding work.
The tests cover areas such as coding deception, broken search, reviewer bypass, warning circumvention, and unauthorized interaction. They deliberately use challenging situations, so they do not measure failure rates in typical use. OpenAI points to the system card for full results.
11. What Do These Models Mean for Developers and Businesses?
Based on OpenAI's reported results, the main change is cost per task. Sol and Luna are shown reaching scores near or above other models on several tests at lower spend, and OpenAI says lower cost and higher usage limits give teams more room to iterate. For agent based applications, the caching controls add another way to manage cost and speed.
These figures are reported by OpenAI. Its evaluations ran in a research environment or through the API, which may differ slightly from production ChatGPT. Competitor results come from publicly available reports, and Fable 5 scores were used where Fable 5.1 scores were unavailable. Teams may want to validate results on their own workloads.
12. API Availability and Access
GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access Luna in the desktop app. Neither model is available in Chat yet.
In the OpenAI API, the model names are gpt-6-sol and gpt-6-luna. Rollout in ChatGPT is gradual, so the models may not appear immediately.
13. Conclusion
GPT-6 Sol and Luna extend the GPT-6 family with two options built around cost efficiency. Astra remains OpenAI's highest-capability choice, Sol is positioned for difficult work tasks, and Luna carries the lowest prices of the three. Reported gains in professional work, coding, computer use, and factuality come with lower prices and better caching.
As adoption grows, independent testing will show how these results hold up in practice. Teams can compare Sol and Luna against their current models on their own tasks, looking at both quality and cost per task.]
