← ModelSwitchModelSwitch · Free Model Test
Test cheaper models against your real work.
Five steps. No live traffic is rerouted — candidate models are shadow-tested against your tasks, and we only recommend a switch when the evidence holds.
Free Model Test$020 tasks3 candidate models
- Current workload
- Workload type
- Tasks
- Evaluation method
- Result
Your current workload
What kind of work is it?
Add your real tasks
Paste tasks, or upload a CSV / JSONL. Each task needs an input. We never send task contents to analytics.
CSV columns: id,input,expectedOutput,metadata. JSONL: one object per line with {"input": "...", "expectedOutput": "..."}.
How we'll evaluate
We pre-select an evaluator from your workload type. You can adjust later.
—
What your report will show
| Metric | Current model | Recommended candidate |
|---|---|---|
| Quality score | measured | measured |
| Baseline quality | measured | — |
| Quality delta | — | measured |
| Cost | measured | measured |
| Cost delta | — | measured |
| Latency | measured | measured |
| Reliability | measured | measured |
| Cost per successful task | measured | measured |
| Decision | — | SAFE_TO_CANARY / MORE_TESTING / DO_NOT_SWITCH |
Decision is never "safe" from token price alone — it combines quality on your tasks, reliability and cost per successful task.