How Much Does the OpenAI API Cost? 3 Real Examples (2026)
2026-09-17
"How much will the OpenAI API actually cost me?" is the first question every founder asks — and the pricing page alone can't answer it, because OpenAI bills by token, not by request. This guide walks through the billing model and three realistic scenarios so you can map them to your own product. For your exact numbers, use our free calculator.
How OpenAI bills you
Every model has three meter rates, charged per 1 million tokens (a token ≈ ¾ of an English word):
- Input tokens — everything you send: system prompt, user message, context, history.
- Cached input tokens — repeated prompt prefixes served from cache, billed at roughly 10% of the input rate.
- Output tokens — everything the model writes back. Always the most expensive meter.
Verified on 2026-09-17 against OpenAI's official standard short-context rates: GPT-6 Astra is $10 input / $50 output per 1M tokens, GPT-5.6 Sol is $4 / $20, Terra is $2 / $12 and Luna is $0.20 / $1.20. Sol's price is promotional through at least 2026-11-21; long-context and Fast mode rates differ.
Example 1: customer-support chatbot
A B2B SaaS adds an AI assistant: 500 monthly active users, 10 questions per user per day, 30 days. That's 150,000 requests. Each carries ~400 input tokens (system prompt + question + a little history) and produces ~300 output tokens. Monthly volume: 60M input + 45M output tokens.
On Terra that's 60 × $2 + 45 × $12 = ~$660/month. On Luna: 60 × $0.20 + 45 × $1.20 = ~$66/month. Same workload, 10× cheaper — which is why quality tests and difficulty-based routing matter more than choosing one model for every request.
Example 2: content generator
A marketing tool writes product descriptions: 20,000 generations/month, short briefs in (~200 tokens), long copy out (~700 tokens). Volume: 4M input + 14M output. On Terra: 4 × $2 + 14 × $12 = ~$176/month. Notice the shape: output is more than 95% of the bill, so capping response length and tightening formats pays immediately.
Example 3: high-volume classifier
An ops team routes 1M tickets/month, ~300 input tokens each, ~10 output tokens (a label). Volume: 300M input + 10M output. On Luna: 300 × $0.20 + 10 × $1.20 = ~$72/month. On Astra the same meter would be ~$3,500. Classification is the clearest case for testing a small model first.
The traps that inflate real invoices
- Prompt growth — teams keep adding instructions; input doubles in months.
- Retries — failed and timed-out calls still consume tokens.
- Context accumulation — chat history grows with every turn unless you trim or summarize it.
- Forgotten environments — dev/staging loops quietly burning tokens.
That's why we recommend budgeting with a 20–30% safety margin — the calculator applies one automatically and shows a low/expected/high range.
Estimate your own cost in two minutes
Plug your users, request frequency and token sizes into our AI API Cost Calculator — it models caching, infrastructure extras and break-even, and lets you compare OpenAI side-by-side with Claude, Gemini and Mistral. Estimates are for planning only; always verify current pricing with OpenAI.