Cost Compare
Compare AI API costs. Monthly pricing across GPT-5.6 Sol/Terra/Luna, Claude Fable 5 / Opus 5 / Sonnet 5, Gemini 3.x, Grok 4.6, and DeepSeek V4.
1,000
100 100K
500
10 100K
10
1 10K
Filter Providers
GPT-4.1 Nano
openai$0.0900/mo
Per Request: $0.0003
Daily: $0.00
Cheapest model in OpenAI lineup. 1M context at $0.10 input.
Gemini 2.5 Flash-Lite
google$0.0900/mo
Per Request: $0.0003
Daily: $0.00
Cheapest model from a Tier-1 provider. Batch: $0.05/$0.20.
Gemini 2.0 Flash
google$0.0900/mo
Per Request: $0.0003
Daily: $0.00
DEPRECATED — shut down June 1, 2026. Migrate to Gemini 2.5 Flash or 3.5 Flash-Lite.
Llama 3.2 90B Vision
meta$0.0900/mo
Per Request: $0.0003
Daily: $0.00
Llama 3.3 70B
meta$0.0990/mo
Per Request: $0.0003
Daily: $0.00
Pricing based on typical Groq/Together AI rates. Can be self-hosted.
GPT-4o Mini
openai$0.1350/mo
Per Request: $0.0004
Daily: $0.00
Legacy. Prefer GPT-5.6 Luna ($0.20 / $1.20) or GPT-5.4 Nano for new work.
Grok 4.1 Fast
x$0.1350/mo
Per Request: $0.0004
Daily: $0.00
2M context is unique at this price tier. Prompt caching included automatically.
DeepSeek V3.2 (Chat)
deepseek$0.1470/mo
Per Request: $0.0005
Daily: $0.00
LEGACY — model string routes to V4 Flash. Do not budget this row for new apps.
Grok 3 Mini
x$0.1650/mo
Per Request: $0.0005
Daily: $0.01
Output at $0.50/M — 4x cheaper than GPT-5 Mini output.
GPT-5.6 Luna
openai$0.2400/mo
Per Request: $0.0008
Daily: $0.01
Listed as gpt-5.6-luna. 1.05M context. Long-context band (prompts over 272K) $0.40 / $1.80 per 1M. Same $0.20 input floor as Grok 4.1 Fast with higher output price.
GPT-5.4 Nano
openai$0.2475/mo
Per Request: $0.0008
Daily: $0.01
Cheapest GPT-5.4 variant. Prefer Luna when you want the current GPT-5.6 cheap tier at the same $0.20 input.
Gemini 3.1 Flash-Lite
google$0.3000/mo
Per Request: $0.0010
Daily: $0.0100
Official: $0.25 / $1.50 for text/image/video. Audio input is $0.50 / 1M.
DeepSeek V4 Flash (Chat)
deepseek$0.3300/mo
Per Request: $0.0011
Daily: $0.0110
Peak (01:00–04:00 and 06:00–10:00 UTC) $0.44 / $1.32 cache-miss. Off-peak is half ($0.22 / $0.66). Cache hit $0.007 off-peak / $0.014 peak. Rates effective Aug 16, 2026. Max output 384K on current V4 Flash.
GPT-4.1 Mini
openai$0.3600/mo
Per Request: $0.0012
Daily: $0.0120
Llama 3.1 405B
meta$0.4050/mo
Per Request: $0.0014
Daily: $0.0135
Gemini 3.5 Flash-Lite
google$0.4650/mo
Per Request: $0.0015
Daily: $0.0155
Official: $0.30 / $2.50 (text/image/video/audio). Batch $0.15 / $1.25.
Gemini 2.5 Flash
google$0.4650/mo
Per Request: $0.0015
Daily: $0.0155
Outstanding value. Configurable reasoning depth. Flat pricing regardless of context.
DeepSeek R1
deepseek$0.4935/mo
Per Request: $0.0016
Daily: $0.0165
Legacy alias. Prefer deepseek-v4-flash thinking. First-party reasoner path was scheduled to route away Jul 24, 2026.
Kimi K2.5
moonshot$0.5550/mo
Per Request: $0.0019
Daily: $0.0185
1T total / 32B active MoE. Thinking and non-thinking modes. 50.2% Humanity's Last Exam. 76% cheaper than Claude Opus 4.5 on comparable tasks. OpenAI-compatible API.
Gemini 3.1 Flash
google$0.5700/mo
Per Request: $0.0019
Daily: $0.0190
Released May 2026. Flash tier of Gemini 3.1 with configurable thinking. Flat pricing regardless of context length.
Grok 4.3
x$0.7500/mo
Per Request: $0.0025
Daily: $0.0250
Official: $1.25 / $0.20 / $2.50 under 200K; $2.50 / $0.40 / $5.00 above. 1M context.
Gemini 3.7 Flash
google$0.7875/mo
Per Request: $0.0026
Daily: $0.0262
Shipped Aug 13, 2026. Introductory $0.75 / $3.75 through 31 Dec 2026, then $1.50 / $7.50 from 1 Jan 2027. Cache $0.075 now / $0.15 in January. Thinking tokens billed as output.
Gemini 3.6 Flash
google$0.7875/mo
Per Request: $0.0026
Daily: $0.0262
Listed as gemini-3.6-flash. Introductory $0.75 / $3.75 through 31 Dec 2026 (matches 3.7 Flash), then $1.50 / $7.50 from 1 Jan 2027. Prefer 3.7 Flash for new work.
Claude 3.5 Haiku
anthropic$0.8400/mo
Per Request: $0.0028
Daily: $0.0280
Retired on first-party Claude API except Bedrock and Google Cloud. Prefer Haiku 4.5.
Kimi K2.6
moonshot$0.8850/mo
Per Request: $0.0029
Daily: $0.0295
Released Apr 20, 2026. API list price checked Aug 2026: $0.95 input / $4.00 output, cache ~$0.16. 256K context. Open weights (Modified MIT).
GPT-5.4 Mini
openai$0.9000/mo
Per Request: $0.0030
Daily: $0.0300
Still listed. Prefer GPT-5.6 Luna ($0.20 / $1.20) when you want the current cheap GPT-5.6 tier.
o4 Mini
openai$0.9900/mo
Per Request: $0.0033
Daily: $0.0330
Replaced o3-mini. Best-value reasoning model in OpenAI lineup.
DeepSeek V4 Pro
deepseek$0.9900/mo
Per Request: $0.0033
Daily: $0.0330
GA as DeepSeek-V4-Pro-0813. Peak cache-miss $1.32 / $3.96; off-peak half ($0.66 / $1.98). Cache hit $0.022 / $0.044. Cost estimator uses peak so budgets are not surprised. Effective Aug 16, 2026.
Claude Haiku 4.5
anthropic$1.05/mo
Per Request: $0.0035
Daily: $0.0350
Budget tier of current Claude generation.
GPT-5.5 Mini
openai$1.08/mo
Per Request: $0.0036
Daily: $0.0360
Released May 2026. Prefer GPT-5.6 Terra ($2 / $12) or Luna ($0.20 / $1.20) for new cheap-volume work.
Gemini 1.5 Pro
google$1.13/mo
Per Request: $0.0037
Daily: $0.0375
Legacy. 2x pricing for prompts >128K tokens. Upgrade to Gemini 2.5 Pro.
Grok 4.6
x$1.50/mo
Per Request: $0.0050
Daily: $0.0500
Official: $2 / $0.50 / $6 per 1M under 200K prompt tokens; $4 / $1 / $12 at or above 200K for the whole request. 500K context. Reasoning effort low/medium/high/xhigh.
Grok 4.5
x$1.50/mo
Per Request: $0.0050
Daily: $0.0500
Official: $2 / $0.30 / $6 under 200K; doubles above. Prefer grok-4.6 for new work.
GPT-4.1
openai$1.80/mo
Per Request: $0.0060
Daily: $0.0600
1M token context at flat rate. Recommended replacement for GPT-4o.
o3
openai$1.80/mo
Per Request: $0.0060
Daily: $0.0600
Reasoning model. Replaced o1 at 87% lower cost with better performance.
Gemini 3.5 Flash
google$1.80/mo
Per Request: $0.0060
Daily: $0.0600
Official: $1.50 / $9.00. Batch/Flex $0.75 / $4.50.
Gemini 2.5 Pro
google$1.88/mo
Per Request: $0.0063
Daily: $0.0625
Pricing doubles above 200K tokens ($2.50/$15). Built-in "thinking" capability.
Claude Sonnet 5
anthropic$2.10/mo
Per Request: $0.0070
Daily: $0.0700
Anthropic made $2 / $10 the standard Sonnet 5 price (the planned Sep 1, 2026 rise to $3 / $15 will not happen). Cache hits $0.20. 4.7+ tokenizer.
GPT-4o
openai$2.25/mo
Per Request: $0.0075
Daily: $0.0750
Legacy model. GPT-5.4 or GPT-4.1 preferred for new projects.
GPT-5.6 Terra
openai$2.40/mo
Per Request: $0.0080
Daily: $0.0800
Listed as gpt-5.6-terra. Long-context band $4 / $18 per 1M. Strong default when Sol is more than you need.
Gemini 3.1 Pro Preview
google$2.40/mo
Per Request: $0.0080
Daily: $0.0800
Official Gemini API: $2 / $12 under 200K prompt tokens, $4 / $18 above. Cache reads $0.20 / $0.40. Thinking tokens billed as output.
GPT-5.4
openai$3.00/mo
Per Request: $0.0100
Daily: $0.1000
Previous flagship. Prefer GPT-5.6 Terra at a similar $2 input band with a lower $12 output rate. 1.05M context.
Claude Sonnet 4.6
anthropic$3.15/mo
Per Request: $0.0105
Daily: $0.1050
Most popular production model. 1M token context at flat rate.
Claude 3.5 Sonnet
anthropic$3.15/mo
Per Request: $0.0105
Daily: $0.1050
Previous generation. Prefer Sonnet 5 ($2 / $10) for new projects. Not on the current first-party price table.
Grok 4
x$3.15/mo
Per Request: $0.0105
Daily: $0.1050
Same pricing as Claude Sonnet 4.6. OpenAI-compatible API format.
Grok 4.1
x$3.15/mo
Per Request: $0.0105
Daily: $0.1050
Released May 2026. Full-size upgrade over Grok 4 with reasoning and prompt caching. OpenAI-compatible API format.
Grok 3
x$3.15/mo
Per Request: $0.0105
Daily: $0.1050
Claude Sonnet 4.7
anthropic$3.47/mo
Per Request: $0.0116
Daily: $0.1155
Released May 2026. Sonnet tier of the Claude 4.7 generation. Uses the same new tokenizer as Opus 4.7 (~35% more tokens than 4.6), so token/cost estimates are adjusted. 1M context flat-rate.
Claude Opus 5
anthropic$5.25/mo
Per Request: $0.0175
Daily: $0.1750
Released Jul 24, 2026. Same $5 / $25 card as prior Opus 4.x. Cache hits $0.50. 4.7+ tokenizer (~30% more tokens vs 4.6).
Claude Opus 4.8
anthropic$5.25/mo
Per Request: $0.0175
Daily: $0.1750
Still on the official price table. Prefer Opus 5 for new projects. 4.7+ tokenizer.
Claude Opus 4.7
anthropic$5.25/mo
Per Request: $0.0175
Daily: $0.1750
Released Apr 16, 2026. 87.6% SWE-bench Verified, 64.3% Terminal-Bench 2.0. New xhigh effort level. 3.75MP vision (3x previous). Task budgets beta. New tokenizer emits up to ~35% more tokens vs Opus 4.6, raising effective cost — token/cost estimates here are adjusted accordingly.
Claude Opus 4.6
anthropic$5.25/mo
Per Request: $0.0175
Daily: $0.1750
Previous flagship. Now superseded by Opus 4.7. Extended thinking support. 1M context flat-rate, no surcharge.
GPT-5.6 Sol
openai$6.00/mo
Per Request: $0.0200
Daily: $0.2000
Listed as gpt-5.6-sol. Standard rates under ~270K input; long context is $10 / $45 per 1M. Batch 50% off. Regional data-residency endpoints add 10%.
GPT-5.5
openai$6.00/mo
Per Request: $0.0200
Daily: $0.2000
Released Apr 24, 2026. Same $5 / $30 card as GPT-5.6 Sol. Cached input $0.50. Long-context band (prompts over 272K) $10 / $45. Prefer Sol or Terra for new work unless you need this snapshot.
Claude Fable 5
anthropic$10.50/mo
Per Request: $0.0350
Daily: $0.3500
Official: $10 / $50 per MTok, cache hits $1. Tokenizer is the 4.7+ generation (~30% more tokens than Sonnet 4.6 and earlier).
Claude Mythos 5
anthropic$10.50/mo
Per Request: $0.0350
Daily: $0.3500
On the official price table at $10 / $50. Limited availability (Glasswing). Prefer Fable 5 if you do not have access.
GPT-5.5 Pro
openai$36.00/mo
Per Request: $0.1200
Daily: $1.20
Premium tier of GPT-5.5. For workloads where accuracy outweighs cost. Cached input listed at 10% of input. Batch/Flex at 50% off.
Features
- Comprehensive matrix comparing 40+ current models across OpenAI, Anthropic, Google, DeepSeek, xAI, Moonshot, Meta, and Mistral
- Dynamic recalculation of monthly SaaS bills based on adjustable Input/Output ratio sliders
- Instant cross-provider scaling: instantly see the financial impact of moving from GPT-5.6 Sol to Luna, Gemini 3.7 Flash, or DeepSeek V4
- Visual indicators for the most cost-effective routing options based on real-time token economics
- Granular filtering to isolate reasoning models, vision models, or ultra-fast sub-second latency models
Common Use Cases
- Auditing a massive cloud AI bill to find exact drop-in replacement models that cut costs by 90%
- Presenting a comparative financial dashboard to executive teams when requesting a monthly generative AI budget
- Developing a Dynamic Model Routing system (LLM Router) that falls back to cheaper APIs for simple classification tasks
- Evaluating whether the price premium of reasoning flagships (Sol, Opus 5, Fable 5) is justified over Luna or Flash
- Calculating the profit margins of an AI wrapper application by modeling cost-per-user per month
Navigating the LLM Price Matrix
The generative AI market is currently segmented into three distinct pricing tiers. Choosing the wrong tier can bankrupt an AI startup overnight.
The Three Tiers of AI Economics:
- Frontier/Reasoning Models (Premium): GPT-5.6 Sol ($5 / $30), Claude Opus 5 ($5 / $25), Claude Fable 5 ($10 / $50). Use these for hard coding, long agents, and work where a miss is expensive.
- Fast/Mini Models (Commodity): GPT-5.6 Luna ($0.20 / $1.20), Gemini 3.7 Flash ($0.75 / $3.75 intro through Dec 2026), Claude Haiku 4.5 ($1 / $5), DeepSeek V4 Flash peak ($0.44 / $1.32). They handle most extraction, classification, and chat.
- Open-Source Local Models (Free Compute): Llama 3.3 70B or Qwen 2.5 72B on your own GPU. You skip per-token API fees and pay for the box instead.
Examples
Valid - Tier 1 Routing (Complex)
Task: Write a full React application.
Model: Claude Sonnet 5 ($2.00 In / $10.00 Out) or GPT-5.6 Terra ($2.00 / $12.00)
Result: Production default, not the cheapest, usually the right first pick. Valid - Tier 2 Routing (Simple)
Task: Extract names from this text into JSON.
Model: GPT-5.6 Luna ($0.20 In / $1.20 Out) or DeepSeek V4 Flash peak ($0.44 / $1.32)
Result: Cheap enough that volume, not the rate card, is the real budget line.Frequently Asked Questions
How much cheaper are "Mini" or "Flash" models compared to the flagship models?
Often by an order of magnitude or more. GPT-5.6 Luna input is $0.20 vs Sol at $5.00 (25x). DeepSeek V4 Flash peak is $0.44 / $1.32 vs Opus 5 at $5 / $25. If you process 100 million input tokens a month, Sol is about $500 and Luna is about $20 before output.
Are open-source models always cheaper?
Not always. While you do not pay per-token API fees for open-source models (like Llama 3) if you host them yourself, you do pay for the GPU server (e.g., $1,000/month for an AWS instance). If your token volume is low, it is actually much cheaper to use a managed API like OpenAI or Anthropic than to rent your own dedicated GPU.
Do any providers offer bulk discounts?
Yes! Both OpenAI and Anthropic offer a "Batch API". If you submit a massive file of requests and are willing to wait up to 24 hours for the results, they will process the tokens at exactly a 50% discount. This is the ultimate hack for offline data processing.
Tips
- Use an "LLM Router" architecture: send all user inputs to a cheap Mini model first. If the Mini model fails or expresses low confidence, only then route the request to the expensive flagship model.
- Pay close attention to "Cached Input" pricing. Providers like Anthropic offer massive 90% discounts if you repeatedly send the exact same long document over and over.