GPT-4.1 vs Gemini 2.5 Pro
Comparing two large-context flagships for long-document processing, retrieval-augmented generation, and code-heavy development work.
Quick answer: Both list 1M-token context windows, but pricing splits by direction: Gemini 2.5 Pro lists the cheaper input ($1.25 vs $2 per 1M) while GPT-4.1 lists the cheaper output ($8 vs $10 per 1M), so which is cheaper depends on your input/output mix.
Pricing & context
All figures below are list prices pulled directly from the Harpd pricing registry (last verified 2026-08-18). Prices change often — open each model’s source link to confirm before budgeting. Machine-readable copy: /data/llm-pricing.json.
| Metric | GPT-4.1 | Gemini 2.5 Pro |
|---|---|---|
| Provider | OpenAI | |
| Input / 1M tokens | $2 | $1 |
| Output / 1M tokens | $8 | $10 |
| Cached input / 1M tokens | $0.2 | $0.13 |
| Batch discount | 50% off | 50% off |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Source | OpenAI pricing ↗ | Google pricing ↗ |
$1,625 vs $1,800 / month (100k calls, 5k in / 1k out)
First benchmark: JSON extraction, 100 real tasks — target ship 2026-09-15
Capabilities
GPT-4.1 pairs a 1M token context window with strong code generation and instruction-following, plus prompt caching to reduce repeat-prefix cost.
Best for
Generation-heavy long-context work — the $8 per 1M output price wins whenever outputs dominate the token mix — and code-centric workflows.
Limitations
At $2 per 1M input it lists above Gemini 2.5 Pro for read-heavy workloads, and there is no documented multimodal parity claim in this comparison.
Capabilities
Gemini 2.5 Pro also provides a 1M token context window, adds native multimodal input and strong reasoning, and includes prompt caching for long-context use. Prompts over 200k tokens bill at higher rates ($2.50 input / $15 output per 1M).
Best for
Input-heavy long-context work, multimodal documents, and retrieval where most tokens are read rather than written.
Limitations
Output at $10 per 1M is the pricier direction of this pair, and the >200k-token tier ($2.50/$15) removes the input-price advantage on very large prompts.
Both models sit at the top of the capability range and share a 1M-token context, so either can handle very long inputs. The better fit depends on whether your workload is code-centric, multimodal, or reasoning-heavy — and on which direction your token mix flows. Price alone will not tell you which is safe to adopt — measure both on your real tasks.
Updated 2026-08-18. Sources: OpenAI and Google official pricing pages (verified 2026-08-18); no Harpd-measured benchmark yet.
Methodology & sources
Prices on this page come from the Harpd pricing registry, which mirrors the officialOpenAI and Google pricing pages and was last verified 2026-08-18. Capability notes summarize documented provider positioning — they are not Harpd measurements. Cheaper-cost claims are computed from the registry at a fixed reference workload, so they are reproducible from the published dataset. Read the pricing methodology and themodel replacement guide before switching a production workload.
Price tells you what a model costs. It does not tell you whether it can replace your current model on your real tasks.
- Benchmarks measured on Harpd are planned — see /benchmarks/.
- Full list-price table across providers: /llm-pricing/.