<aside> 💰 TL;DR: AI API prices have dropped 60-90% since 2024. DeepSeek V3.2 is the cheapest capable model at $0.28/M input. Gemini 2.0 Flash-Lite is the absolute cheapest at $0.075/M. Use Crazyrouter to access all of them through one API key at even lower prices.
</aside>
The AI API market in 2026 is a race to the bottom — and developers are winning. Since GPT-4 launched at $30/M input tokens in 2023, prices have plummeted. Today, you can get GPT-4-level performance for under $0.30/M tokens. But with 20+ major models across 6 providers, choosing the right one for your budget is harder than ever.
This guide breaks down every major model's pricing, helps you pick the best value for your use case, and shows you how to save even more with an API gateway.
| Model | Input/M | Output/M | Context | Best For |
|---|---|---|---|---|
| GPT-5.2 Pro | $21.00 | $168.00 | 200K | Hardest reasoning |
| GPT-5.2 | $1.75 | $14.00 | 200K | Coding & agents |
| GPT-5 | $1.25 | $10.00 | 128K | General flagship |
| GPT-5 Mini | $0.25 | $2.00 | 200K | Best value |
| GPT-5 Nano | $0.05 | $0.40 | 128K | High throughput |
| o3 | $2.00 | $8.00 | 200K | Mid-tier reasoning |
| o4-mini | $1.10 | $4.40 | 200K | Reasoning value |
💡 OpenAI Batch API gives 50% off all models for async processing (results within 24h). Cached input tokens are 90% cheaper.
| Model | Input/M | Output/M | Context | Best For |
|---|---|---|---|---|
| Claude Opus 4.6 | $5.00 | $25.00 | 200K | Complex analysis |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 200K | Coding & all-round |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | Fast classification |
💡 Opus 4.6 is 67% cheaper than Opus 4.1. Prompt caching saves 90% on input tokens. Batch API adds another 50% off. Claude 3 Haiku ($0.25/$1.25) retires April 2026.
| Model | Input/M | Output/M | Context | Best For |
|---|---|---|---|---|
| Gemini 3.1 Pro | $2.00-$4.00 | $12.00-$18.00 | 200K+ | Next-gen flagship |
| Gemini 2.5 Pro | $1.25 | $10.00 | 2M | Long documents |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | Fast mid-tier |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | Cheapest mainstream |
| Gemini 2.0 Flash-Lite | $0.075 | $0.30 | — | Absolute cheapest |
💡 Gemini 2.5 Pro has the largest context window at 2M tokens — 10x larger than GPT-5 or Claude. Most models have a free tier for prototyping.
| Model | Input/M | Output/M | Notes |
|---|---|---|---|
| DeepSeek V3.2 | $0.28 | $0.42 | Best value, 90% cache discount |
| DeepSeek R1 | $0.55 | $2.19 | Reasoning model, open source |
| Llama 4 | Free | Free | Self-host, open source |
💡 DeepSeek V3.2 delivers GPT-4-class performance at 1/7th the price of GPT-5. Open-source models like Llama 4 are free to self-host but require GPU infrastructure.