AI model pricing FAQ

The questions every team asks before committing AI spend — answered plainly, with the math.

prices verified September 26, 2026

Tokens and tasks

What's the difference between token pricing and task pricing?

Token pricing is the provider's rate card: dollars per million input tokens and per million output tokens. It's what tables like the JetAI dashboard show. Task pricing is what you actually pay to complete one unit of work — tokens consumed × their rates.

The two diverge constantly. A model with half the token price costs more per task if it uses twice as many tokens, generates verbose output, or needs two retries where a stronger model succeeds once. This is why our coding-model ranking uses cost per task, not cost per token.

What counts as an input token vs an output token?

Input tokens are everything you send: the system prompt, conversation history, documents, tool definitions, and the user's message. Output tokens are everything the model generates — including its thinking tokens on reasoning models. That last part surprises people: a "cheap" reasoning model that thinks for 5,000 tokens before answering bills all of it at output rates.

Why is output always more expensive than input?

Generating tokens is computationally harder than reading them. Input tokens are processed in parallel in one pass; output tokens must be generated one at a time, each depending on the last. The typical price ratio is 3–5x — GPT-6 Sol charges $2/M in and $10/M out; Claude Opus 5.5 charges $4/M in and $20/M out. Consequence: verbosity is the biggest cost driver. Cap max tokens and ask for concise formats.

Prompt caching

How does prompt caching reduce my bill?

When you resend context the model has already processed — system prompts, tool schemas, documents, conversation history — providers skip reprocessing and bill those tokens at a cache-read rate instead of the fresh-input rate. The discounts are dramatic:

  • Claude Opus 5.5: $0.20/M cache reads vs $4.00/M fresh — 95% off
  • Grok 4.7: $0.50/M cached input vs $2.00/M fresh — 75% off
  • DeepSeek: cache hits at roughly 2% of the input rate

Cache-hit rates above 50% are common in production agent loops, where each step re-sends the whole conversation. Design prompts with stable content first and the changing query last — caches key on the prompt prefix, so stable-prefix structure earns the discount.

Is prompt caching automatic?

Mostly. Anthropic and xAI cache eligible prefixes automatically; OpenAI discounts cached input on the GPT-6 family automatically too. You don't flip a switch — you design for it: keep system prompts byte-identical across calls, don't shuffle document order, and put volatile content at the end of the prompt.

Promotional vs list prices

What's the difference between promotional and list prices?

List prices are the standard rate card. Promotional prices are time-limited discounts designed to pull you onto a model. Two live examples (September 2026):

  • GPT-5.6 Sol at $4/$20 — promotional pricing through at least November 21, 2026
  • Gemini 3.8 Flash at $0.75/$3.75 — introductory pricing through December 31, 2026, then it doubles to $1.50/$7.50

The trap: building your unit economics on a promo rate, then discovering your margins evaporate when it expires. Always budget against the price the model reverts to, and treat the promo period as a bonus, not the baseline.

How do batch discounts work?

OpenAI, Google, and Anthropic discount async batch requests roughly 50% when you don't need an instant answer — e.g., GPT-6 Sol drops from $2/$10 to $1/$5 on Batch/Flex. Nightly jobs, bulk classification, dataset labeling, and report generation should never pay interactive rates.

Estimating monthly spend

How do I estimate my monthly AI spend?

Five steps:

  1. Measure tokens per task. Run a sample of your workload and record average input and output tokens. Don't guess — tokenizers differ across models.
  2. Multiply by volume. Tasks per month × tokens per task = monthly token totals, split into input and output.
  3. Apply rates. (input_M × input_price) + (output_M × output_price), using current list prices from the dashboard.
  4. Apply your cache-hit rate. Estimate what share of input is repeated context and price it at the cache-read rate.
  5. Add a buffer. 25–50% for retries, longer-than-average outputs, and traffic growth.

Or skip the spreadsheet: the dashboard's model picker runs steps 2–4 for you from live prices.

Worked example: 100K support chats a month
Per chat: 2K input (0.5K cached) + 0.5K output
Model: DeepSeek V4.1 Flash at $0.15/$0.60 off-peak (cache hits $0.003/M)
Monthly input: 200M → (150M × $0.15 + 50M × $0.003) ≈ $22.65
Monthly output: 50M × $0.60 = $30.00
Total: ≈ $52.65/month for 100K chats

That's the power of cheap input tiers on volume workloads: a hundred thousand conversations for the price of a nice dinner. Run the same math on a frontier model and watch it land tens to nearly 100x higher.

What hidden costs don't price lists show?
  • Thinking tokens on reasoning models, billed at output rates
  • Long-context price cliffs — xAI doubles the whole request's rate past 200K prompt tokens; Google's pricing page likewise steps Pro-tier rates up past 200K prompts. Check the current thresholds before routing long-context workloads.
  • Tokenizer differences — the same text tokenizes differently per model, so billed tokens per task vary. Run a sample of your own prompts through each API and compare billed tokens rather than assuming parity.
  • Platform markups — Bedrock, Vertex AI, and Azure add their own margins over direct API pricing
  • Peak windows — DeepSeek charges $0.30/$1.20 at peak vs $0.15/$0.60 off-peak
  • Taxes and currency on the final invoice