The three tiers at a glance
All figures below are USD per 1 million tokens (input / output), verified against OpenAI's official API pricing page on September 26, 2026 and cross-checked on the JetAI pricing dashboard.
| Model | Input / 1M | Output / 1M | Role |
|---|---|---|---|
| GPT-6 AstraBatch/Flex: $5/$25 · long context: $20/$75 | $10.00 | $50.00 | Flagship reasoning |
| GPT-6 SolBatch/Flex: $1/$5 · fast mode: $4/$20 · long context: $4/$15 | $2.00 | $10.00 | Balanced default |
| GPT-6 LunaBatch/Flex: $0.05/$0.25 | $0.10 | $0.50 | Budget volume |
The naming maps neatly to use cases: Astra is the star — maximum capability, maximum price. Sol is the workhorse sun most workloads orbit. Luna is the cheap moon you put bulk jobs on.
GPT-6 Astra — $10 / $50: when the flagship earns its keep
Astra is OpenAI's top reasoning tier, aimed at work where a wrong answer costs real money: complex multi-step agentic workflows, hard mathematical reasoning, scientific analysis, and long-horizon planning tasks. At $10 input and $50 output per million tokens, it is priced 5x above Sol — so it only makes sense when cheaper tiers measurably fail.
A crucial detail: thinking models bill their reasoning at output rates. Astra's chain-of-thought tokens cost $50/M just like final answer tokens. A long-thinking Astra call can burn several thousand output tokens before it even starts answering. Always model thinking-token volume when budgeting, not just final-answer length.
Cost-reduction paths exist: the Batch/Flex API drops Astra to $5/$25 if your workload tolerates async processing, and prompt caching discounts repeated context. If you run Astra in production, those two levers typically cut the effective bill 30–60%.
GPT-6 Sol — $2 / $10: the default for most teams
Sol is the tier most production deployments should start on. Five times cheaper than Astra, it handles customer chat, document Q&A, drafting, and mid-complexity coding work at what is usually indistinguishable-from-flagship quality on routine tasks.
The two variants worth knowing:
- Batch/Flex mode — $1/$5. Half price for non-urgent jobs: nightly report generation, bulk classification, dataset labeling. If nothing needs an answer in seconds, there is no reason to pay the interactive rate.
- Fast mode — $4/$20. A lower-latency configuration that costs double. Pay it only where latency directly affects conversion or user retention, not by default.
At $2/$10, a million-token document corpus costs $2 to read once and $10 per million tokens of generated output. For reference, that makes Sol roughly comparable in price to Grok 4.7 ($2/$6) while sitting below Claude Opus 5.5 ($4/$20).
GPT-6 Luna — $0.10 / $0.50: the bulk tier
Luna is OpenAI's cheapest current-generation tier — a full order of magnitude below Sol. It exists for one job: enormous volumes of simple work. Classification, sentiment tagging, summarization at scale, first-draft marketing copy, and data-extraction pipelines all fit.
The economics are striking: one million input tokens cost ten cents. You can run Luna over a hundred million tokens of input for $10. With Batch/Flex, that halves again to $0.05/$0.25.
The trade-off is capability. Luna is not the model for reasoning-heavy coding agents or judgment calls. The pattern that works: Luna does the first pass at scale, and only uncertain or high-value items escalate to Sol or Astra. That router architecture — cheap first pass, expensive escalation — is the single most effective cost structure in production AI, and Luna is built for the cheap end of it.
Cost per task: what the bill actually looks like
Token prices are the sticker; task prices are the bill. Here's the same job — a 50K-token document plus a 2K-token summary — across all three tiers:
Sol: 0.05 × $2.00 + 0.002 × $10.00 = $0.12 per task
Astra: 0.05 × $10.00 + 0.002 × $50.00 = $0.60 per task
Same task, 100x spread between Luna and Astra. Run 10,000 of those summaries a month and you're looking at $60 on Luna versus $6,000 on Astra. That is why the first question is always "what's the cheapest tier that handles this task reliably?" — not "what's the best model?"
Which tier should you pick?
Pick Astra when…
- The task genuinely needs frontier reasoning and cheaper tiers have been measured failing
- A wrong answer costs more than the price difference (compliance, finance, scientific work)
- You're running a small number of high-stakes decisions, not bulk volume
Pick Sol when…
- You're building production features and don't yet know your failure profile — it's the sane default
- You need a balance of quality and cost across mixed workloads (chat, docs, coding)
- You can batch non-urgent work and use the $1/$5 Flex rate
Pick Luna when…
- Volume is the story: millions of tokens of classification, extraction, or summarization
- Tasks are simple enough that quality differences don't matter
- You're building a routing architecture with Luna as the first-pass tier
And if even Luna feels expensive for your simplest jobs, compare it against the cheapest coding models — DeepSeek V4.1 Flash and Grok 4.7 compete in the same budget band with different strengths.
Quick answers
How much does GPT-6 Astra cost?
As of September 2026, GPT-6 Astra is listed at $10 per million input tokens and $50 per million output tokens. Batch/Flex processing brings that to $5/$25, and long-context requests cost $20/$75.
Is GPT-6 Sol the best value in the family?
For most production workloads, yes. At $2/$10 per million tokens — 5x cheaper than Astra — Sol delivers near-flagship quality on routine tasks. It is the default pick unless you need Astra's frontier reasoning or Luna's bulk pricing.
When should I use GPT-6 Luna?
Use Luna ($0.10/$0.50 per million tokens) for high-volume, low-complexity work: classification, summarization at scale, draft copy, and data extraction where quality demands are modest.
Do GPT-6 prices include prompt caching discounts?
Listed prices are for fresh tokens. OpenAI discounts cached input tokens on the GPT-6 family, which can substantially reduce bills for workloads with repeated system prompts or document context.