Which Model, at Which Effort? A Two-Mode Routing Card
- Figure it out
- Claude Opus 5 high. Real consequence: Claude Opus 5 max.
- Execute it
- Kimi K3 max. Scale/value: DeepSeek V4 Pro 0813 max.
Choose one:
- Figure it out when choosing the right solution is the hard part and a plausible mistake could survive your checks.
- Execute it when the target and method are settled and checks can establish whether the result is correct.
Writing code is usually Execute — but only once the implementation is decided. If your tests can verify the code and still not tell you whether the architecture, the diagnosis, or the plan was right, that part is Figure.
Get the default for your setup
The answers above assume an any-model harness. Pick Claude Code or Codex when your harness limits which models you can run. Open the advanced controls only when your model access or acceptable spend differs.
Adjust budget or model access13 of 13 models selected
Maximum model spend for one task on this benchmark workload—not an invoice or cost per successful real task.
Maximum model spend for one task on this benchmark workload—not an invoice or cost per successful real task.
Models you can use
Is the hard part choosing the right solution, or carrying out one already defined?
Figure it out
Choose the solution when judgment is the hard part and a plausible wrong answer could survive the available checks.
Execute it
Carry out a settled target and method when checks can establish whether the result is correct.
Compare the measured frontiers
Higher is better; left is cheaper. One frontier per mode: Figure on graded work-product analysis, Execute on four execution benchmarks averaged.
Measured frontiers
Higher is better; left is cheaper. One frontier per mode, each scored and priced on the same benchmarks.
How the defaults are picked
Each mode chooses the highest measured quality within one benchmark-task spend cap. If nothing fits, the card shows the cheapest compatible result as over budget.
| Mode | Measured on | Default cap |
|---|---|---|
| Figure it out | AA-Briefcase Analytical Quality Elo | $12.00 |
| Execute it | GDPval, Terminal-Bench 2.1, tau3-Banking, SciCode | $1.00 |
Scale/value Execute lowers the cap to $0.10. Real-consequence Figure removes the cap. Both use the same selector, not a separate model rule.
What each mode is measured on
Execute it averages four public benchmarks, each counted once: GDPval, Terminal-Bench 2.1, tau3-Banking and SciCode. The same four set the price — each evaluation's own cost per task, averaged the same way — so the score and the bill describe one task. Every measured configuration on this card carries all four.
That mix is more expensive than the one the card used to price execution work on, which was Terminal-Bench and SciCode alone — short tasks. GDPval tasks are full work products, so a dollar of benchmark spend covers less ground than it looks like it should. The $1.00 cap admits 21 of the 30 measured configurations.
Figure it out uses AA-Briefcase Analytical Quality Elo, which grades open-ended professional work products rather than checkable answers, priced on that board's own cost per task. It is the only public board measuring the thing this mode is for, and its coverage is the honest cost: Briefcase has scored 14 of the 30 configurations. The other 16 are absent from this mode rather than approximated.
Neither mode uses Artificial Analysis's Coding or Agentic Index. Those are AA's own weighted summaries of the boards below them, and the Coding Index was withdrawn in August 2026 — it went from scoring every configuration here to scoring none, in a week, while every board underneath it stayed live. Summaries can be retired; measurements are what a recommendation should rest on.
Data and limitssource, raw rows, and caveats
- Artificial Analysis supplies every score and every benchmark-task cost. The two mode scores are computed here from those published boards, not read off AA's Coding or Agentic Index. Vendor documentation supplies pricing and effort semantics.
- Gemini 3.7 Flash shipped on August 13, but it is not routed yet. Google's public Gemini API catalog captured for this update still stops at Gemini 3.6 Flash, and no independently scored Artificial Analysis 3.7 row was discoverable. The card therefore keeps Gemini 3.6 Flash as the measured Google Flash row and does not reuse its scores for 3.7. This will settle when Google's 3.7 API contract and a comparable AA row are both published.
- The August 13 release is DeepSeek V4 Pro 0813, not a new Flash checkpoint:
deepseek-v4-flashstill points to V4 Flash 0731. AA scores Pro 0813 max on all four Execute benchmarks, putting it on that frontier at $0.064 per execute task. - AA has not published a Briefcase row for Pro 0813, so Figure excludes it rather than inferring judgment quality from execution evidence. Pro 0813 high, low, and non-thinking are also unmeasured; the card does not project them.
- DeepSeek documents
low / high / maxthinking effort, with thinking on by default. The capture uses AA's observed August 13 task cost; DeepSeek's different Pro tariff beginning August 16 will be rechecked after it reaches AA rather than applied early. - xAI exposes low, medium, high, and xhigh effort under one
grok-4.6API ID; AA has measured only high. The other three settings remain named gaps. xAI's $2/$0.50/$6 headline pricing applies only below 200K prompt tokens; at 200K or more, the whole request uses $4/$1/$12 input/cached/output rates. - Benchmark-task spend is not your invoice. It excludes retries, review time, latency, subscriptions, and the cost of a wrong answer.
- Deduced rows remain visible as dated research bets but cannot become recommendations.
| Model · effort | AA Index | Execute | $ / execute task | $ / knowledge task | Evidence |
|---|---|---|---|---|---|
| Claude Opus 5 · max | 63.05 | 63.58 | $2.8238 | $0.2982 | printed |
| Claude Opus 5 · xhigh | 62.52 | 63.04 | $2.1610 | $0.2252 | printed |
| Claude Fable 5 · max | 62.07 | 61.23 | $3.6782 | $0.4465 | printed |
| Claude Opus 5 · high | 61.48 | 62.10 | $1.4518 | $0.1536 | printed |
| Claude Fable 5 · xhigh | 61.34 | — | — | — | deduced |
| GPT-5.6 Sol · max | 60.93 | 62.43 | $1.4925 | $0.2339 | printed |
| Grok 4.6 · high | 60.92 | 63.76 | $1.1601 | $0.0590 | printed |
| Kimi K3 · max | 59.70 | 62.20 | $0.9910 | $0.2408 | printed |
| Claude Fable 5 · high | 59.43 | — | — | — | deduced |
| GPT-5.6 Sol · xhigh | 59.01 | 60.68 | $1.0189 | $0.1159 | printed |
| Claude Opus 5 · medium | 58.64 | 57.87 | $0.8509 | $0.0821 | printed |
| Claude Fable 5 · medium | 57.84 | — | — | — | deduced |
| GPT-5.6 Sol · high | 57.33 | 59.26 | $0.6957 | $0.0676 | printed |
| Claude Opus 4.8 · max | 57.33 | 56.64 | $2.5607 | $0.4015 | printed |
| GPT-5.6 Terra · max | 56.58 | 59.01 | $0.6159 | $0.1338 | printed |
| Grok 4.5 · high | 55.76 | 57.27 | $0.5021 | $0.0539 | printed |
| GPT-5.6 Sol · medium | 55.57 | 57.92 | $0.4871 | $0.0364 | printed |
| Claude Sonnet 5 · max | 55.26 | 56.58 | $1.9641 | $0.4932 | printed |
| Claude Opus 4.8 · xhigh | 54.57 | — | — | — | deduced |
| Claude Fable 5 · low | 53.79 | — | — | — | deduced |
| DeepSeek V4 Pro 0813 · max | 53.00 | 55.48 | $0.0642 | $0.0158 | printed |
| GPT-5.6 Terra · xhigh | 52.77 | 53.76 | $0.3783 | $0.0499 | printed |
| Claude Opus 5 · low | 52.46 | 50.63 | $0.5116 | $0.0281 | printed |
| GPT-5.6 Luna · max | 52.32 | 54.65 | $0.0536 | $0.0181 | printed |
| Claude Opus 4.8 · high | 52.08 | — | — | — | deduced |
| DeepSeek V4 Flash 0731 · max | 51.77 | 55.21 | $0.0333 | $0.0059 | printed |
| Gemini 3.6 Flash · high | 51.58 | 51.55 | $0.7677 | $0.0466 | printed |
| Claude Sonnet 5 · xhigh | 51.03 | — | — | — | deduced |
| GPT-5.6 Sol · low | 50.73 | 52.10 | $0.2991 | $0.0198 | printed |
| GPT-5.6 Terra · high | 50.11 | 51.24 | $0.2738 | $0.0323 | printed |
| GPT-5.6 Luna · xhigh | 50.06 | 52.00 | $0.0368 | $0.0099 | printed |
| Claude Opus 4.8 · medium | 50.06 | — | — | — | deduced |
| Claude Sonnet 5 · high | 48.88 | — | — | — | deduced |
| Kimi K3 · low | 48.25 | 53.42 | $0.3560 | $0.0141 | printed |
| GPT-5.6 Luna · high | 46.96 | 48.45 | $0.0262 | $0.0055 | printed |
| GPT-5.6 Terra · medium | 46.76 | 48.20 | $0.1525 | $0.0139 | printed |
| Claude Opus 4.8 · low | 44.96 | — | — | — | deduced · extrapolated |
| Claude Sonnet 5 · medium | 42.58 | — | — | — | deduced · extrapolated |
| GPT-5.6 Terra · low | 41.30 | 42.07 | $0.1226 | $0.0080 | printed |
| GPT-5.6 Luna · medium | 38.91 | 38.88 | $0.0143 | $0.0018 | printed |
| GPT-5.6 Luna · low | 33.85 | 33.66 | $0.0113 | $0.0009 | printed |
| Claude Sonnet 5 · low | 32.61 | — | — | — | deduced · extrapolated |
Use the normal route when you want one answer quickly. Change the setup only when your available models or budget actually differ.