Methodology · Cost per successful task

Token price is not the cost

A model that is cheap per token but fails often costs more once you pay for retries and throw away failed runs. The number that matters is cost per task that actually succeeded.

Quick answer

What is cost per successful task?

Cost per successful task is total workload cost divided by successful tasks. It folds in retries, failed runs and wasted tokens, so a cheap-per-token model that fails often is usually more expensive than a pricier model that passes the first time. The per-task failure rates Harpd shows are illustrative until the measured benchmark ships.

Formula
totalCost ÷ successfulTasks
Counts
retries + failures + waste
Status
Rates illustrative

Data source: Harpd cost methodology

Why it matters

Per-token price hides the cost that actually shows up on a bill — retries and failed runs. Cost per successful task is the metric production systems pay against, so publishing the definition and the denominator keeps cost comparisons honest instead of token-theoretic.

Limitations

  • The per-task failure rates in the examples are illustrative, not measured — Harpd’s benchmark that would produce real rates is PLANNED / MODELED (target ship 2026-09-15).
  • Your real cost depends on your own observed pass rates, caching, batching and provider discounts, not the illustrative figures.
  • Cost per successful task compares models on the same task; it does not by itself judge quality, only cost efficiency given a pass rate.

Data source: Harpd cost methodology

Answer

Cost per successful task = total workload cost ÷ successful tasks. It folds in retries, failed runs, and wasted tokens. A cheap-per-token model that fails 30% of the time is usually more expensive than a pricier model that passes the first time.

Evidence
  • Accounts for retries and failed runs
  • Surfaces false economies of flaky models
  • Comparable across models on the same task

The definition

// Cost per successful taskcostPerSuccessfulTask = totalWorkloadCost / successfulTasks// where totalWorkloadCost includes every attempt:totalWorkloadCost = sum(cost of each attempt) // incl. retries & failuressuccessfulTasks = count(tasks whose final attempt passed)

Why cheap-per-token is misleading

Token price only describes one attempt. In production, a task may need several attempts before it succeeds — or it may fail entirely and be discarded. The cheap model's low per-token rate is multiplied by more attempts, while a more capable (and more expensive) model may finish in one or two tries.

Consider a task where a cheap model passes 70% of the time and needs an average of 1.4 attempts, versus a pricier model that passes 98% of the time in 1.02 attempts. Divide the higher per-token price by the pass rate and the "expensive" model is often cheaper per real outcome.

What goes into the denominator

  • Retries. Automatic re-attempts when the first output fails a validation or quality gate.
  • Failures. Tasks that never pass and are dropped — their tokens are spent and gone.
  • Wasted output. Partial or malformed completions that are thrown away before a correct one is produced.

All of these are included in totalWorkloadCost. Dividing by successfulTasks — never by attempts — is what makes the metric honest about what you actually pay for.

Honesty note

The per-task failure rates above are illustrative, not measured. Harpd's benchmark that would produce real measured rates is in a PLANNED / MODELED stage (target ship 2026-09-15). Until then, use the calculator with your own observed pass rates, or plug in modeled rates clearly labeled as such.