Cheap model + low success rate = expensive outcome.
Most "AI cost" pages quote the per-call price. That number is wrong for production agents. The real bill is cost per successful task — what you actually pay for an outcome, not a request. Estimate yours above, and learn why the metric matters below.
Estimate your AI bill, model by model.
Cheap model + low success rate = expensive outcome. Estimate cost per task that actually shipped — not per API call.
| Model | $ / M input | $ / M output | Monthly | vs. current |
|---|---|---|---|---|
| Llama 3.1 8BMeta · Together | $0.18 | $0.18 | $108 | −96% |
| GPT-4o miniOpenAI | $0.15 | $0.60 | $135 | −96% |
| Qwen 2.5 72BAlibaba · OpenRouter | $0.40 | $0.40 | $240 | −92% |
| DeepSeek V3DeepSeek | $0.27 | $1.10 | $245 | −92% |
| Gemini 2.5 FlashGoogle | $0.30 | $2.50 | $400 | −87% |
| Llama 3.1 70BMeta · Together | $0.88 | $0.88 | $528 | −82% |
| Claude Haiku 4Anthropic | $0.80 | $4.00 | $800 | −73% |
| Mistral Large 2Mistral | $2.00 | $6.00 | $1,600 | −47% |
| Gemini 2.5 ProGoogle | $1.25 | $10.00 | $1,625 | −46% |
| GPT-4.1OpenAI | $2.00 | $8.00 | $1,800 | −40% |
| Claude Sonnet 4Anthropic | $3.00 | $15.00 | $3,000 | current |
| GPT-5OpenAI | $5.00 | $20.00 | $4,500 | +50% |
| Claude Opus 4Anthropic | $15.00 | $75.00 | $15,000 | +400% |
Prices reflect typical 2026 list pricing and may change. Verify with each provider before you commit to a budget.
Price alone doesn't tell you if a cheaper model can safely replace yours.
See the live benchmark data →Three numbers to keep separate.
Cost per call
What every pricing page quotes. Useful for capacity planning; almost useless for budgeting an agent. A $0.001 call that fails 30% of the time and triggers a $0.05 retry is not a $0.001 call.
Cost per attempt
Cost per call × average attempts per successful task. Captures retries. Still missing the real cost of partial answers that need human review.
Cost per successful task
What every production system should budget against. The full spend, divided by the count of tasks that actually shipped a usable answer. The flagship metric for Harpd ModelSwitch.
Two models, same workload, very different bill.
| Claude Sonnet 4 | DeepSeek V3 | |
|---|---|---|
| List price (5k in / 1k out) | $450 / month | $57 / month |
| Success rate on 200 real tasks | 94% | 78% |
| Retries needed (avg) | 1.06 | 1.28 |
| Effective cost per successful task | $5.06 | $5.84 |
| Verdict | Wins on quality | Cheaper per call, more expensive per outcome |
Hypothetical example. Your success rate, retry pattern and prompt shape will be different. ModelSwitch measures these numbers on your real tasks.
Cost per successful task, answered
What is "cost per successful task"?
monthly spend ÷ number of tasks that actually shipped a usable result. Two models can have the same per-call cost and wildly different cost per successful task because their success rates differ.