Free developer tool · No signup

LLM API Cost Calculator

Compare token pricing across GPT, Claude, Gemini, DeepSeek and Llama, and estimate what your chatbot, RAG app or agent will actually cost per request, per day and per month. Real, source-linked pricing - last verified August 2026.

Start from a common workload

Prompt + context sent to the model

Tokens the model generates back

Per day (billed ×30)

Show total per:
≈ 1 token ≈ 0.75 words · 1,000 words ≈ 1,300 tokens
Lowest cost for this workload
Gemini 1.5 Flash $20.25/month
Up to $4,592/month cheaper than Claude 3 Opus
ModelTierIn $/1MOut $/1MCost / requestTotal / month
Gemini 1.5 Flashcheapest
Google
Efficient$0.075$0.3$0.00014
$20.25
GPT-4.1 nano
OpenAI
Efficient$0.1$0.4$0.00018
$27.00
Gemini 2.0 Flash
Google
Efficient$0.1$0.4$0.00018
$27.00
Llama 3.3 70B
Meta (hosted) · Typical hosted rate
Open / Value$0.13$0.4$0.00020
$30.60
GPT-4o mini
OpenAI
Efficient$0.15$0.6$0.00027
$40.50
DeepSeek V3
DeepSeek
Open / Value$0.27$1.1$0.00049
$73.65
Claude 3 Haiku
Anthropic
Efficient$0.25$1.25$0.00051
$76.88
GPT-4.1 mini
OpenAI
Balanced$0.4$1.6$0.00072
$108
Llama 3.1 405B
Meta (hosted) · Typical hosted rate
Open / Value$0.9$0.9$0.00095
$142
DeepSeek R1
DeepSeek
Reasoning$0.55$2.19$0.00099
$148
Claude 3.5 Haiku
Anthropic
Efficient$0.8$4$0.00164
$246
o4-mini
OpenAI
Reasoning$1.1$4.4$0.00198
$297
Gemini 2.5 Pro
Google · ≤200K context tier
Flagship$1.25$10$0.00350
$525
GPT-4.1
OpenAI
Flagship$2$8$0.00360
$540
o3
OpenAI
Reasoning$2$8$0.00360
$540
GPT-4o
OpenAI
Flagship$2.5$10$0.00450
$675
Claude Sonnet 4
Anthropic
Flagship$3$15$0.00615
$923
Claude 3 Opus
Anthropic
Flagship$15$75$0.0307
$4,613

Prices are standard published per-token rates last verified August 2026 and may vary by region, cached/batch tiers and volume commitments. Always confirm on the official pricing page before budgeting a production workload. Llama rates reflect typical hosted-provider pricing. Reasoning models may bill hidden "thinking" tokens as output.

How LLM API pricing works

Every major LLM API bills per token - not per request. A token is a chunk of text roughly ¾ of a word, so about 1,000 English words ≈ 1,300 tokens. Providers quote two separate prices, usually per 1 million tokens:

  • Input tokens - everything you send: the system prompt, chat history, and any retrieved context (RAG).
  • Output tokens - everything the model generates back. These typically cost 3–5× more than input tokens because each output token needs its own forward pass.
cost = (in_tokens / 1e6 × in_price)
    + (out_tokens / 1e6 × out_price)
total = cost × requests

The practical takeaway: output length usually drives your bill. Trimming verbose responses, capping max_tokens, caching repeated context, and routing easy requests to a cheaper model are the four highest-leverage ways to cut LLM cost - often by 5–10× - without hurting quality.

Frequently asked questions

Which LLM API is cheapest?

For most general workloads the lowest-cost options are small or open models - Gemini 1.5/2.0 Flash, GPT-4o mini, Claude 3 Haiku and DeepSeek V3. Flagships like GPT-4.1, Claude Sonnet 4 and Gemini 2.5 Pro cost more but reason better. The right pick depends on your input-vs-output token mix, which this calculator makes obvious.

Why are output tokens more expensive?

Reading your prompt is a single pass, but generating a response runs the model once per output token. That extra compute is why most providers charge 3–5× more for output than input tokens - and why response length is usually the biggest cost driver.

How accurate are these numbers?

They are standard published rates last verified in August 2026, linked to each provider's official pricing page. Pricing changes often and can vary by region, cached/batch tiers and volume commitments, so treat these as planning estimates and confirm on the official page for production budgeting.

Do reasoning models cost more than the table shows?

Often yes. Reasoning models (o3, o4-mini, DeepSeek R1) generate hidden 'thinking' tokens that are billed as output. A request that returns 500 visible tokens might bill several thousand. Add headroom to output tokens when estimating reasoning-model cost.

Learn to build cost-efficient AI systems

Estimating cost is step one. In the Thrive With AI professional bootcamp you build real LLM apps - RAG assistants, agents, production APIs - and learn the token-budgeting, caching and model-routing patterns that keep them cheap at scale.