LLM API Cost Calculator
Compare token pricing across GPT, Claude, Gemini, DeepSeek and Llama, and estimate what your chatbot, RAG app or agent will actually cost per request, per day and per month. Real, source-linked pricing - last verified August 2026.
Start from a common workload
Prompt + context sent to the model
Tokens the model generates back
Per day (billed ×30)
| Model | Tier | In $/1M | Out $/1M | Cost / request | Total / month |
|---|---|---|---|---|---|
Gemini 1.5 Flashcheapest Google | Efficient | $0.075 | $0.3 | $0.00014 | $20.25 |
GPT-4.1 nano OpenAI | Efficient | $0.1 | $0.4 | $0.00018 | $27.00 |
Gemini 2.0 Flash Google | Efficient | $0.1 | $0.4 | $0.00018 | $27.00 |
Llama 3.3 70B Meta (hosted) · Typical hosted rate | Open / Value | $0.13 | $0.4 | $0.00020 | $30.60 |
GPT-4o mini OpenAI | Efficient | $0.15 | $0.6 | $0.00027 | $40.50 |
DeepSeek V3 DeepSeek | Open / Value | $0.27 | $1.1 | $0.00049 | $73.65 |
Claude 3 Haiku Anthropic | Efficient | $0.25 | $1.25 | $0.00051 | $76.88 |
GPT-4.1 mini OpenAI | Balanced | $0.4 | $1.6 | $0.00072 | $108 |
Llama 3.1 405B Meta (hosted) · Typical hosted rate | Open / Value | $0.9 | $0.9 | $0.00095 | $142 |
DeepSeek R1 DeepSeek | Reasoning | $0.55 | $2.19 | $0.00099 | $148 |
Claude 3.5 Haiku Anthropic | Efficient | $0.8 | $4 | $0.00164 | $246 |
o4-mini OpenAI | Reasoning | $1.1 | $4.4 | $0.00198 | $297 |
Gemini 2.5 Pro Google · ≤200K context tier | Flagship | $1.25 | $10 | $0.00350 | $525 |
GPT-4.1 OpenAI | Flagship | $2 | $8 | $0.00360 | $540 |
o3 OpenAI | Reasoning | $2 | $8 | $0.00360 | $540 |
GPT-4o OpenAI | Flagship | $2.5 | $10 | $0.00450 | $675 |
Claude Sonnet 4 Anthropic | Flagship | $3 | $15 | $0.00615 | $923 |
Claude 3 Opus Anthropic | Flagship | $15 | $75 | $0.0307 | $4,613 |
Prices are standard published per-token rates last verified August 2026 and may vary by region, cached/batch tiers and volume commitments. Always confirm on the official pricing page before budgeting a production workload. Llama rates reflect typical hosted-provider pricing. Reasoning models may bill hidden "thinking" tokens as output.
How LLM API pricing works
Every major LLM API bills per token - not per request. A token is a chunk of text roughly ¾ of a word, so about 1,000 English words ≈ 1,300 tokens. Providers quote two separate prices, usually per 1 million tokens:
- Input tokens - everything you send: the system prompt, chat history, and any retrieved context (RAG).
- Output tokens - everything the model generates back. These typically cost 3–5× more than input tokens because each output token needs its own forward pass.
+ (out_tokens / 1e6 × out_price)
total = cost × requests
The practical takeaway: output length usually drives your bill. Trimming verbose responses, capping max_tokens, caching repeated context, and routing easy requests to a cheaper model are the four highest-leverage ways to cut LLM cost - often by 5–10× - without hurting quality.
Official pricing sources
Frequently asked questions
Which LLM API is cheapest?
For most general workloads the lowest-cost options are small or open models - Gemini 1.5/2.0 Flash, GPT-4o mini, Claude 3 Haiku and DeepSeek V3. Flagships like GPT-4.1, Claude Sonnet 4 and Gemini 2.5 Pro cost more but reason better. The right pick depends on your input-vs-output token mix, which this calculator makes obvious.
Why are output tokens more expensive?
Reading your prompt is a single pass, but generating a response runs the model once per output token. That extra compute is why most providers charge 3–5× more for output than input tokens - and why response length is usually the biggest cost driver.
How accurate are these numbers?
They are standard published rates last verified in August 2026, linked to each provider's official pricing page. Pricing changes often and can vary by region, cached/batch tiers and volume commitments, so treat these as planning estimates and confirm on the official page for production budgeting.
Do reasoning models cost more than the table shows?
Often yes. Reasoning models (o3, o4-mini, DeepSeek R1) generate hidden 'thinking' tokens that are billed as output. A request that returns 500 visible tokens might bill several thousand. Add headroom to output tokens when estimating reasoning-model cost.
Learn to build cost-efficient AI systems
Estimating cost is step one. In the Thrive With AI professional bootcamp you build real LLM apps - RAG assistants, agents, production APIs - and learn the token-budgeting, caching and model-routing patterns that keep them cheap at scale.