Measure on real tasks, publish the method
Every number Harpd shows traces back to a documented method. We measure on real tasks, with published methods, and we separate measured results from modeled estimates — always.
Quick answer
What is Harpd's methodology?
Harpd measures AI cost and model quality on real tasks, with every method written down and reproducible. We use four methodologies — verified model pricing, safe model replacement, cost per successful task, and reproducible benchmarks — and we always separate measured results from modeled estimates.
- Methodologies
- 4 documented methods
- Principle
- Real tasks, published methods
- Estimates
- Labeled, never mixed
Data source: Harpd methodology
Why it matters
Documented, reproducible methods are what let a reader or an AI system trust a number. By publishing the exact procedure and separating measured from modeled results, Harpd turns each figure into something anyone can verify or replay — the foundation of citation-grade data.
Limitations
- Methodology pages describe the procedure; the numbers themselves live on the data and research pages and may lag the latest run.
- Benchmarks are planned: not every model family has a completed, published run yet.
- Model prices are sourced from official provider pages and can change between verification timestamps; always check the data date.
Data source: Harpd methodology
Harpd uses four methodologies: sourcing verified model prices, deciding when a cheaper model is safe to switch to, computing cost per successful task, and running reproducible real-task benchmarks. Each is written down and reproducible.
- Verified prices from official provider pages
- Real-task quality gates for model replacement
- Cost per successful task, not per token
- Reproducible benchmark runs with published artifacts
The four methodologies
Each links to a full, detail-level write-up.
1 · Model pricing
How we source every model price from official provider pages, attach a verification timestamp and source URL, flag deprecated models, and publish a single machine-readable source of truth.
Live data2 · Model replacement
How we use shadow testing and per-task quality gates to decide when a cheaper model is actually safe to switch to — and the rule that we only switch when the candidate passes your bar.
3 · Cost per successful task
The metric that matters: total workload cost divided by successful tasks. Cheap-per-token models that fail often are not actually cheap once retries are counted.
4 · Benchmarks
How we run reproducible, real-task evaluations: same prompt across multiple runs, published success definitions, and CSV/JSON artifacts so anyone can replay the result.
PlannedOur measurement principle
We hold one line across all four methodologies:
We measure on real tasks, with published methods, and we separate measured results from modeled estimates.
Three consequences follow from that principle:
- Real tasks, not toy prompts. A model is judged on the work you actually run — JSON extraction, code generation, classification, prose — not a generic leaderboard question.
- Published methods. Every methodology page states the exact procedure, the success definition, and the data used, so the measurement is repeatable.
- Measured vs. modeled, never mixed. If a figure is modeled or not yet measured, it is labeled as such and visually distinguished. We never present an estimate as a result.