Claude vs GPT vs Gemini: API Pricing Compared (2026)
2026-09-17
Comparing LLM providers on price is harder than it looks: each one meters input, output and cache differently, and a "cheap" model can lose on a real workload. Here we run one identical workload through current standard price lists, verified 2026-09-17 on each provider's official page.
The test workload
A typical assistant feature: 150,000 requests/month, ~400 input tokens and ~300 output tokens per request → 60M input + 45M output tokens monthly. No caching, no extras. Estimated token bill per model:
- Gemini 2.5 Flash ($0.30 in / $2.50 out): ~$131/month
- Gemini 2.5 Flash-Lite ($0.10 / $0.40): ~$24/month
- GPT-5.6 Luna ($0.20 / $1.20): ~$66/month
- Claude Haiku 4.5 ($1 / $5): ~$285/month
- Claude Sonnet 5 ($2 / $10): ~$570/month
- Gemini 3.1 Pro Preview ($2 / $12): ~$660/month
- GPT-5.6 Terra ($2 / $12): ~$660/month
More than 25× separates the cheapest and priciest entries in this list — which is why model choice, not provider loyalty, is the biggest cost decision. Run your own workload in the calculator to see your ranking, then compare Terra vs Sonnet 5 and Luna vs Haiku 4.5.
Caching changes the ranking
If your requests share a long stable prefix (system prompt, tool definitions, knowledge snippets), cached-input pricing matters as much as the sticker price:
- OpenAI: cached input at ~10% of the input rate (e.g. Terra: $0.20 vs $2), applied to eligible repeated prefixes.
- Anthropic: cache reads at 10% of input, but cache writes cost 1.25× input — caching pays off after a single reuse within the 5-minute window.
- Google Gemini: explicit context caching at roughly 10% of input, with storage-time billing on some tiers.
An input-heavy RAG app with a 60–80% cache-hit rate can flip the comparison table entirely. Set the cache slider in the calculator to your expected hit rate and watch the ranking change.
When each provider makes sense
- Claude (Anthropic) — long-context work, agent workflows and careful writing; Sonnet 5's $2/$10 launch rate is now its standard rate.
- GPT (OpenAI) — a broad model ladder; Luna, Terra, Sol and Astra support explicit routing by workload difficulty.
- Gemini (Google) — the Flash tiers are among the cheapest capable models anywhere, and integrate naturally with a Google Cloud stack.
The honest answer
There is no universally cheapest provider — there's a cheapest provider for your token shape (input-heavy vs output-heavy, cacheable vs not). Model it in two minutes with the AI API Cost Calculator, and re-check quarterly because providers reprice frequently. Estimates are for planning only — always verify current prices on each provider's official page.