Methodology · Free calculator

Cheap model + low success rate = expensive outcome.

Most "AI cost" pages quote the per-call price. That number is wrong for production agents. The real bill is cost per successful task — what you actually pay for an outcome, not a request. Estimate yours above, and learn why the metric matters below.

Cost per successful task

Estimate your AI bill, model by model.

Cheap model + low success rate = expensive outcome. Estimate cost per task that actually shipped — not per API call.

Estimated % of calls that ship a usable answer. Calibrate from your own data.
Current monthly cost$65
Cheapest alternative$65−99%
Cost per successful task$0at 92% success
Model$ / M input$ / M outputMonthlyvs. current
GPT-5 nanoOpenAI$0.05$0.40$65−99%
Mistral Small 3.1Mistral$0.10$0.30$80−98%
Gemini 2.5 Flash-LiteGoogle$0.10$0.40$90−98%
Qwen 2.5 72BAlibaba · OpenRouter$0.40$0.40$240−95%
DeepSeek V3DeepSeek$0.27$1.10$245−95%
GPT-5 miniOpenAI$0.25$2.00$325−94%
Gemini 2.5 FlashGoogle$0.30$2.50$400−92%
Llama 3.1 70BMeta · Together$0.88$0.88$528−89%
Claude Haiku 4.5Anthropic$1.00$5.00$1,000−80%
Mistral Large 3Mistral$2.00$6.00$1,600−68%
GPT-5OpenAI$1.25$10.00$1,625−68%
Gemini 2.5 ProGoogle$1.25$10.00$1,625−68%
GPT-4.1OpenAI$2.00$8.00$1,800−64%
Claude Sonnet 5Anthropic$2.00$10.00$2,000−60%
Claude Sonnet 4.5Anthropic$3.00$15.00$3,000−40%
Claude Opus 4.8Anthropic$5.00$25.00$5,000+0%

Prices reflect typical 2026 list pricing and may change. Verify with each provider before you commit to a budget.

Price alone doesn't tell you if a cheaper model can safely replace yours.

See the live benchmark data →
Why this metric

Three numbers to keep separate.

A

Cost per call

What every pricing page quotes. Useful for capacity planning; almost useless for budgeting an agent. A $0.001 call that fails 30% of the time and triggers a $0.05 retry is not a $0.001 call.

B

Cost per attempt

Cost per call × average attempts per successful task. Captures retries. Still missing the real cost of partial answers that need human review.

C

Cost per successful task

What every production system should budget against. The full spend, divided by the count of tasks that actually shipped a usable answer.

Worked example

Two models, same workload, very different bill.

Claude Sonnet 4DeepSeek V3
List price (5k in / 1k out)$450 / month$57 / month
Success rate on 200 real tasks94%78%
Retries needed (avg)1.061.28
Effective cost per successful task$5.06$5.84
VerdictWins on qualityCheaper per call, more expensive per outcome

Hypothetical example. Your success rate, retry pattern and prompt shape will be different. Measure these numbers on your real tasks before making a production change.

Cost per successful task, answered

What is "cost per successful task"?
It is the metric that matters when you are running an agent or a production LLM pipeline: monthly spend ÷ number of tasks that actually shipped a usable result. Two models can have the same per-call cost and wildly different cost per successful task because their success rates differ.
Why is cost per successful task different from cost per call?
A $0.001 call that fails 40% of the time is effectively $0.00167 per call that ships, before you account for retries. Add retries, and the real cost multiplies again. Cost per successful task folds success rate and retry cost into one number.
What is a good success rate?
For production agents: 90%+ on the tasks you actually ship. For one-off chat: not really applicable. For RAG extraction: 95%+ is the bar. The default 92% in the calculator is a reasonable planning estimate; calibrate from your own data.
How do I measure cost per successful task on my real workload?
Run a representative set of real tasks against candidate models, record the full cost, and divide it by the number of outputs that meet your success definition. The free AI Cost Calculator above estimates the same metric using your assumed success rate.
Is the cheapest model always the best cost per successful task?
No. A 10× cheaper model that succeeds on 50% of tasks is the same cost per successful task as the flagship, before you account for retries and downstream rework. The methodology page explains why.
How should I measure model success?
Prepare 20–200 real tasks with reference answers or a judge prompt, and run them against each candidate model. Decide pass/fail with the same rubric for every model, then compute cost per successful task from the measured pass rate and total spend.