Free calculator · Updated 2026

How much is your AI workload actually costing you?

Start here for the big picture: pick the model you run today, plug in your call volume and average token mix, and see your estimated monthly bill. For a pure per-model API comparison, open the LLM cost calculator; for autonomous agents, the agent calculator; for two models side by side, the model comparison.

Free calculator

Estimate your AI bill, model by model.

See what your current AI workload actually costs — and which model could do it cheaper without changing the result.

Current monthly cost$65
Cheapest alternative$65−99%
Model$ / M input$ / M outputMonthlyvs. current
GPT-5 nanoOpenAI$0.05$0.40$65−99%
Mistral Small 3.1Mistral$0.10$0.30$80−98%
Gemini 2.5 Flash-LiteGoogle$0.10$0.40$90−98%
Qwen 2.5 72BAlibaba · OpenRouter$0.40$0.40$240−95%
DeepSeek V3DeepSeek$0.27$1.10$245−95%
GPT-5 miniOpenAI$0.25$2.00$325−94%
Gemini 2.5 FlashGoogle$0.30$2.50$400−92%
Llama 3.1 70BMeta · Together$0.88$0.88$528−89%
Claude Haiku 4.5Anthropic$1.00$5.00$1,000−80%
Mistral Large 3Mistral$2.00$6.00$1,600−68%
GPT-5OpenAI$1.25$10.00$1,625−68%
Gemini 2.5 ProGoogle$1.25$10.00$1,625−68%
GPT-4.1OpenAI$2.00$8.00$1,800−64%
Claude Sonnet 5Anthropic$2.00$10.00$2,000−60%
Claude Sonnet 4.5Anthropic$3.00$15.00$3,000−40%
Claude Opus 4.8Anthropic$5.00$25.00$5,000+0%

Prices reflect typical 2026 list pricing and may change. Verify with each provider before you commit to a budget.

Price alone doesn't tell you if a cheaper model can safely replace yours.

Compare model costs →

How the estimate is calculated

monthly cost = calls × input tokens × ($input / 1M) + calls × output tokens × ($output / 1M)

Worked example. At the default reference workload of 100,000 calls per month (5,000 input / 1,000 output tokens), Claude Sonnet 4.5 costs$3,000 per month at list price. The calculator applies the same formula to every model in the registry and ranks the results cheapest first.

Limitations

  • The estimate is a planning estimate, not a bill: prompt caching, batching, free tiers, image token multipliers and negotiated enterprise rates are not modeled.
  • Prices are published list prices in USD per 1M tokens; DeepSeek peak/off-peak differences are shown at the off-peak rate with a caveat.
  • A cheaper model is not a better deal if it fails more tasks — compare cost per successful task via the cost-per-successful-task methodology.
  • The calculator runs entirely in your browser; no workload is sent to or logged by Harpd.

Pricing source & data

List prices come from each provider's official pricing page and were last verified 2026-08-18. The same registry powers the LLM pricing table, themachine-readable dataset (/data/llm-pricing.json) and theHarpd evidence register. How the prices are collected and kept current:model pricing methodology.

How the calculator works

One formula, every model.

01

Pick your model

Start with whatever you are paying for today — Claude Sonnet 4 is the default because it is the most common "expensive baseline."

02

Enter the workload

How many calls per month, average input tokens per call, average output tokens per call. Don't overthink the numbers — directionally correct is enough.

03

Read the table

Every model is ranked cheapest first. The "current" badge marks your baseline; the percentage is what each alternative would save you.

04

Verify before switching

Cheap is not the same as safe. Run your real tasks against the top 3 cheaper models before you commit. Follow the model evaluation guide.

AI cost calculator, answered

How accurate is the AI cost calculator?
It uses 2026 list pricing for the major hosted models and a simple per-token model: (calls × input tokens × $input/M) + (calls × output tokens × $output/M). Your actual bill may differ because of prompt caching, batching, free tiers, or negotiated enterprise rates. The calculator is a planning estimate, not an invoice.
Which model is the default in the calculator?
Claude Sonnet 4 — the most common "expensive default" people compare against. Switch the dropdown to your real model and the table re-ranks immediately. The cheapest model in the table is highlighted; the percentage saving is computed against your current model.
Why is my real OpenAI bill different from this?
OpenAI charges for cached input tokens at a lower rate, gives image tokens a different multiplier, and offers batch API discounts. The calculator ignores all of that and uses a single list price for clarity. If you have detailed billing, paste your last invoice into the same formula and you will see the gap.
Does the calculator send my data anywhere?
No. The whole estimator runs in your browser — there is no network request. The four inputs (model, calls, input tokens, output tokens) never leave the page. We do not log your workload.
What is the cheapest LLM in 2026?
For open-weight models on hosted inference, Llama 3.1 8B (via Together) and Qwen 2.5 72B (via OpenRouter) are typically the cheapest. For hosted proprietary models, GPT-4o mini and Gemini 2.5 Flash dominate the budget tier. Lower prices do not guarantee equivalent results on hard reasoning tasks; test candidates against your own acceptance criteria.
What is the next step after I see a cheaper model?
Run your own 20–200 real tasks against it. Price is one axis; quality on your workload is the other. Compare outputs with reference answers or a defined evaluation rubric before changing production models.