A cheaper model is only safe if it passes
Cheaper-per-token is not the same as cheaper-in-practice. We decide a switch with shadow testing, the same prompt across runs, and a quality gate that fits the task type.
Quick answer
How does Harpd decide a cheaper model is safe to switch to?
Harpd shadow-tests a candidate model on your real tasks with the same prompt across runs, then applies a per-task quality gate — JSON schema checks, unit tests, accuracy/F1, or an LLM judge plus human sample. The switch happens only when the candidate passes your quality bar; a lower price is never, by itself, a reason to switch.
- Method
- Shadow test + quality gate
- Gates
- per task type
- Trigger
- candidate passes your bar
Data source: Harpd model replacement methodology
Why it matters
A model that is cheaper per token can quietly cost more once failures are counted — yet switching blindly to save a cent is how teams ship regressions. Publishing the exact gate per task type lets a reader or an AI system audit whether a 'safe switch' claim is actually backed by evaluation, not price alone.
Limitations
- The quality gates are defined per task type; a candidate that passes one task type is not automatically safe for another.
- Shadow testing observes without serving, so it validates the candidate’s behavior but cannot predict every edge case in live traffic.
- The decision to switch stays with your quality threshold; Harpd reports the pass rate and cost difference, it does not auto-switch.
Data source: Harpd model replacement methodology
We shadow-test a candidate model on your real tasks using the exact same prompt across runs, then apply a per-task quality gate: JSON schema checks, unit tests for code, accuracy/F1 for classification, and an LLM judge plus human sample for prose. We switch only when the candidate passes your quality bar.
- Shadow test on real tasks, not toy prompts
- Same prompt, multiple runs for stability
- Quality gate varies by task type
- Switch only when the candidate passes your bar
The procedure
- 1
Shadow test
The candidate model runs in parallel with your current model on live or replayed traffic. Production is never affected — the candidate is observed, not serving.
- 2
Same prompt, multiple runs
We hold the prompt constant and run the candidate several times to measure stability, not a lucky single shot. Variance matters as much as the average.
- 3
Real-task evaluation
Outputs are scored against the task's own success definition — the same one your production code would use — not a generic benchmark score.
- 4
Quality gate per task type
Each task type has its own gate (see below). If the candidate clears the gate at a lower cost, it is a candidate for switch.
Quality gates by task type
| Task type | Quality gate |
|---|---|
| Structured extraction | JSON schema validation — output must parse and satisfy the schema |
| Code generation | Unit tests must pass against the generated code |
| Classification | Accuracy / F1 against a labeled set above your threshold |
| Prose / open-ended | LLM judge plus a human-sampled review for tone, factuality, and instruction-following |
The rule
Switch only when the candidate passes your quality bar. A lower price is never, by itself, a reason to switch.
The gate is yours to set. Harpd reports the candidate's pass rate and the cost difference; the decision to switch stays with your quality threshold. Nothing is switched automatically without a gate passing.