LLM API Cost Calculator
Enter how many tokens each request uses and how many requests you expect, and compare what the same workload would cost across popular models from OpenAI, Anthropic, Google and DeepSeek.
Prompt + system prompt + context
The model's reply
Share of input served from the prompt cache
Monthly volume
30,000 requests
75,000,000 tokens
Cost comparison
| Model | Input / 1M | Output / 1M | Per request | Per day | Per month |
|---|---|---|---|---|---|
| GPT-6 LunaOpenAI | $0.1 | $0.5 | $0.00045 | $0.45 | $13.50 |
| GPT-5.6 LunaOpenAI | $0.2 | $1.2 | $0.00100 | $1.00 | $30.00 |
| GPT-5.4 nanoOpenAI | $0.2 | $1.25 | $0.00103 | $1.03 | $30.75 |
| DeepSeek FlashDeepSeek · Peak-hour price; off-peak is 50% less | $0.3 | $1.2 | $0.00120 | $1.20 | $36.00 |
| Gemini 3.5 Flash-LiteGoogle | $0.3 | $2.5 | $0.00185 | $1.85 | $55.50 |
| Gemini 2.5 FlashGoogle | $0.3 | $2.5 | $0.00185 | $1.85 | $55.50 |
| Gemini 3.8 FlashGoogle · Promotional price through Dec 31, 2026; $1.50 / $7.50 from Jan 1, 2027 | $0.75 | $3.75 | $0.00337 | $3.38 | $101 |
| GPT-5.4 miniOpenAI | $0.75 | $4.5 | $0.00375 | $3.75 | $113 |
| Claude Haiku 4.5Anthropic | $1 | $5 | $0.00450 | $4.50 | $135 |
| DeepSeek V4 ProDeepSeek · Peak-hour price; off-peak is 50% less | $1.32 | $3.96 | $0.00462 | $4.62 | $139 |
| Gemini 3.5 FlashGoogle | $1.5 | $9 | $0.00750 | $7.50 | $225 |
| Gemini 2.5 ProGoogle | $1.25 | $10 | $0.00750 | $7.50 | $225 |
| GPT-6 SolOpenAI | $2 | $10 | $0.00900 | $9.00 | $270 |
| Claude Sonnet 5Anthropic | $2 | $10 | $0.00900 | $9.00 | $270 |
| GPT-5.6 TerraOpenAI | $2 | $12 | $0.01 | $10.00 | $300 |
| Gemini 3.1 Pro (preview)Google | $2 | $12 | $0.01 | $10.00 | $300 |
| Claude Sonnet 4.6Anthropic | $3 | $15 | $0.01 | $13.50 | $405 |
| GPT-5.6 SolOpenAI | $4 | $20 | $0.02 | $18.00 | $540 |
| Claude Opus 5.5Anthropic | $4 | $20 | $0.02 | $18.00 | $540 |
| Claude Opus 5Anthropic | $5 | $25 | $0.02 | $22.50 | $675 |
| GPT-5.5OpenAI · Price for prompts under 272K tokens | $5 | $30 | $0.03 | $25.00 | $750 |
| GPT-6 AstraOpenAI | $10 | $50 | $0.05 | $45.00 | $1,350 |
| Claude Fable 5.1Anthropic | $10 | $50 | $0.05 | $45.00 | $1,350 |
Standard (pay-as-you-go) API prices in USD. Batch APIs, off-peak discounts and enterprise agreements can lower costs; image, audio, tool and web search charges aren't included. Different tokenizers turn the same text into different token counts, so compare models with your own measured usage.
Data last updated:
Results are estimates for planning only. Prices, limits and real-world usage vary — always confirm against the provider's official documentation or your own measurements before making decisions.
Compare LLM API costs before you build
The same AI feature can cost a few dollars or several thousand a month depending on which model you pick. Prices vary by more than a hundred times between the smallest and largest models, and output tokens usually cost several times more than input tokens. This calculator puts popular models from OpenAI, Anthropic, Google and DeepSeek side by side for your exact workload, so you can budget realistically and see where a smaller model could save money.
How to use the calculator
- Enter the average input tokens per request — system prompt, user message and any documents or chat history.
- Enter the average output tokens per request — the length of the model’s reply.
- If part of your prompt repeats across requests, set the percentage that will be served from the prompt cache.
- Enter requests per day and working days per month.
- Filter providers, sort the table and compare cost per request, per day and per month. The cheapest option is highlighted.
Not sure how many tokens your prompts use? Paste a real example into the LLM Token Counter first.
How the cost is calculated
Each provider charges a price per million input tokens and a price per million output tokens. For one request, the cost is (input tokens × input price + output tokens × output price) ÷ 1,000,000. Cached input tokens are charged at the provider’s lower cache price where one is published. Daily and monthly costs multiply that by your request volume.
Some models charge more for very long prompts. Gemini 3.1 Pro and Gemini 2.5 Pro, for example, switch to higher rates when the prompt exceeds 200,000 tokens; the calculator applies those long-context rates automatically. OpenAI also lists long-context rates, but its pricing page doesn’t state where they begin, so standard rates are used for OpenAI models. DeepSeek prices shown are peak-hour rates; off-peak usage costs half as much.
Where the prices come from
All prices are standard pay-as-you-go API rates in US dollars, taken from each provider’s official pricing page and last checked on 2026-09-27. Batch processing (often 50% off), promotional pricing, enterprise discounts, and charges for images, audio, web search or tools aren’t included. Prices change frequently, so always confirm on the provider’s pricing page before you commit.
Ways to lower your LLM bill
- Right-size the model: route simple tasks such as classification or extraction to a small, cheap model.
- Cap output length: ask for concise answers and set a max tokens limit.
- Use prompt caching: keep long, unchanging instructions at the start of the prompt.
- Batch non-urgent jobs: most providers discount asynchronous batch requests.
- Trim context: retrieve only the relevant passages instead of whole documents.
Thinking about self-hosting an open model instead? Estimate the hardware with the GPU VRAM Calculator, and use the Percentage Calculator to compare savings between options.
Frequently asked questions
How is LLM API cost calculated?
Providers charge separately for input tokens (what you send) and output tokens (what the model writes), priced per million tokens. Cost per request = input tokens × input price + output tokens × output price, divided by 1,000,000.
Why are output tokens more expensive?
Generating text is slower and more compute-intensive than reading it, because the model produces output one token at a time. Output prices are typically three to eight times the input price.
What is cached input?
If many requests start with the same text — a long system prompt or document — providers can cache it and charge much less for the repeated part. Enter the share of your input that's cached to see the savings for models with published cache prices.
Are these prices up to date?
They're taken from each provider's official pricing page on the date shown on this page, for standard pay-as-you-go usage. Providers change prices often, so check the official page before committing to a budget.
Is the cheapest model the best choice?
Not necessarily. A more capable model may need shorter prompts, fewer retries or less human review. Test a few models on your real task, then use the calculator to compare costs at your volume.
Related tools
LLM Token Counter
Count tokens for OpenAI models exactly, estimate Claude and Gemini tokens, and see the cost.
GPU VRAM Calculator for LLMs
Estimate the GPU memory needed to run or fine-tune an LLM, and which GPUs fit.
Percentage Calculator
Work out percentages, percentage change, increases and decreases.