Start with the job, inspect the evidence, then price your own workload. Recommendations compare models within one benchmark source and evaluation setup. They cover the evidence we have, not every model on the market.
What are you building?
What matters most?
Writing, refactoring, or debugging code. Quality measured by the LiveBench coding average (code generation and completion on questions refreshed each release). Models LiveBench has not scored fall back to Aider polyglot, which stopped publishing in October 2025, compared only with other Aider results. Cost is dollars for a 225-case benchmark run: Aider's measured spend where it ran the model, otherwise projected from list prices.
Quality threshold: 20.0% on livebench-coding (fallback when it ranks nothing: aider-polyglot). Ranking cost uses measured Aider-run cost when available, otherwise a token-usage projection. Your workload estimate below uses your own token counts.
Picks from comparable evidence
Comparison group: livebench · livebench-coding · release=2026-06-25; score=mean of task means per category. 32 evidence groups tracked; 53 records excluded from this recommendation because of comparison coverage, pricing, or the quality threshold.
For models without a recorded benchmark-run cost: Estimated default: 10,600 input and 4,800 output tokens per completed case, across 225 assumed cases. These are fallback assumptions, not a measured run or your workload estimate.
Estimated default: 10,600 input and 4,800 output tokens per completed case, across 225 assumed cases. These are fallback assumptions, not a measured run or your workload estimate.
livebench-coding · release=2026-06-25; score=mean of task means per category
Estimated default: 10,600 input and 4,800 output tokens per completed case, across 225 assumed cases. These are fallback assumptions, not a measured run or your workload estimate.
livebench-coding · release=2026-06-25; score=mean of task means per category
Estimated default: 10,600 input and 4,800 output tokens per completed case, across 225 assumed cases. These are fallback assumptions, not a measured run or your workload estimate.
livebench-coding · release=2026-06-25; score=mean of task means per category
Verified purchase destination unavailable for this provider.
What will my workload cost?
Compare exact models and providers before you fill up.
Updating estimates for the selected workload…
glm-5.3-flash [release=2026-06-25; score=mean of task means per category]
baidu
Loading estimate…
deepseek-v4-flash [release=2026-06-25; score=mean of task means per category]
venice
Loading estimate…
Estimates use tracked list rates. They are not actual charges or a billing commitment. Taxes, tool fees, negotiated rates and provider-specific limits may differ. Provider sign-in is required to buy credits.
All contenders shown belong to the comparison group above. Benchmark settings and price observations can change; inspect source evidence before relying on a result.