Jul 2026· updated Aug 2026

Which Model, at Which Effort? A Two-Mode Routing Card

TL;DR
Figure it out
Claude Opus 5 high. Real consequence: Claude Opus 5 max.
Execute it
Kimi K3 max. Scale/value: DeepSeek V4 Pro 0813 max.

Choose one:

  • Figure it out when choosing the right solution is the hard part and a plausible mistake could survive your checks.
  • Execute it when the target and method are settled and checks can establish whether the result is correct.

Writing code is usually Execute — but only once the implementation is decided. If your tests can verify the code and still not tell you whether the architecture, the diagnosis, or the plan was right, that part is Figure.

Get the default for your setup

The answers above assume an any-model harness. Pick Claude Code or Codex when your harness limits which models you can run. Open the advanced controls only when your model access or acceptable spend differs.

Your setup
Adjust budget or model access13 of 13 models selected
$

Maximum model spend for one task on this benchmark workload—not an invoice or cost per successful real task.

$

Maximum model spend for one task on this benchmark workload—not an invoice or cost per successful real task.

Models you can use

Is the hard part choosing the right solution, or carrying out one already defined?

Figure it out

Choose the solution when judgment is the hard part and a plausible wrong answer could survive the available checks.

Claude Opus 5· high$10.41 / Briefcase task$12.00 cap
Real consequence → Claude Opus 5 max $17.79 / Briefcase taskJust above cap → Claude Opus 5 xhigh $14.26 / Briefcase task

Execute it

Carry out a settled target and method when checks can establish whether the result is correct.

Kimi K3· max$0.99 / execute task$1.00 cap
Scale/value → DeepSeek V4 Pro 0813 max $0.06 / execute taskJust above cap → Grok 4.6 high $1.16 / execute task

Compare the measured frontiers

Higher is better; left is cheaper. One frontier per mode: Figure on graded work-product analysis, Execute on four execution benchmarks averaged.

Measured frontiers

Higher is better; left is cheaper. One frontier per mode, each scored and priced on the same benchmarks.

6861056142517952165$0.10$0.50$2.0$10Briefcase cost per task (USD, log scale)AA-Briefcase Analytical Quality EloDEFAULTPREMIUMCONSEQUENCE
OpenAI
Sol
Anthropic
Opus 5Fable 5Opus 4.8Sonnet 5
Moonshot
Kimi K3
DeepSeek
DeepSeek V4 Flash 0731
xAI
Grok 4.6Grok 4.5
Google
3.6 Flash
Dashed line = measured Pareto frontier · flags = default, presets, and point just above the cap · data: August 13, 2026

How the defaults are picked

Each mode chooses the highest measured quality within one benchmark-task spend cap. If nothing fits, the card shows the cheapest compatible result as over budget.

ModeMeasured onDefault cap
Figure it outAA-Briefcase Analytical Quality Elo$12.00
Execute itGDPval, Terminal-Bench 2.1, tau3-Banking, SciCode$1.00

Scale/value Execute lowers the cap to $0.10. Real-consequence Figure removes the cap. Both use the same selector, not a separate model rule.

What each mode is measured on

Execute it averages four public benchmarks, each counted once: GDPval, Terminal-Bench 2.1, tau3-Banking and SciCode. The same four set the price — each evaluation's own cost per task, averaged the same way — so the score and the bill describe one task. Every measured configuration on this card carries all four.

That mix is more expensive than the one the card used to price execution work on, which was Terminal-Bench and SciCode alone — short tasks. GDPval tasks are full work products, so a dollar of benchmark spend covers less ground than it looks like it should. The $1.00 cap admits 21 of the 30 measured configurations.

Figure it out uses AA-Briefcase Analytical Quality Elo, which grades open-ended professional work products rather than checkable answers, priced on that board's own cost per task. It is the only public board measuring the thing this mode is for, and its coverage is the honest cost: Briefcase has scored 14 of the 30 configurations. The other 16 are absent from this mode rather than approximated.

Neither mode uses Artificial Analysis's Coding or Agentic Index. Those are AA's own weighted summaries of the boards below them, and the Coding Index was withdrawn in August 2026 — it went from scoring every configuration here to scoring none, in a week, while every board underneath it stayed live. Summaries can be retired; measurements are what a recommendation should rest on.

Data and limitssource, raw rows, and caveats
  • Artificial Analysis supplies every score and every benchmark-task cost. The two mode scores are computed here from those published boards, not read off AA's Coding or Agentic Index. Vendor documentation supplies pricing and effort semantics.
  • Gemini 3.7 Flash shipped on August 13, but it is not routed yet. Google's public Gemini API catalog captured for this update still stops at Gemini 3.6 Flash, and no independently scored Artificial Analysis 3.7 row was discoverable. The card therefore keeps Gemini 3.6 Flash as the measured Google Flash row and does not reuse its scores for 3.7. This will settle when Google's 3.7 API contract and a comparable AA row are both published.
  • The August 13 release is DeepSeek V4 Pro 0813, not a new Flash checkpoint: deepseek-v4-flash still points to V4 Flash 0731. AA scores Pro 0813 max on all four Execute benchmarks, putting it on that frontier at $0.064 per execute task.
  • AA has not published a Briefcase row for Pro 0813, so Figure excludes it rather than inferring judgment quality from execution evidence. Pro 0813 high, low, and non-thinking are also unmeasured; the card does not project them.
  • DeepSeek documents low / high / max thinking effort, with thinking on by default. The capture uses AA's observed August 13 task cost; DeepSeek's different Pro tariff beginning August 16 will be rechecked after it reaches AA rather than applied early.
  • xAI exposes low, medium, high, and xhigh effort under one grok-4.6 API ID; AA has measured only high. The other three settings remain named gaps. xAI's $2/$0.50/$6 headline pricing applies only below 200K prompt tokens; at 200K or more, the whole request uses $4/$1/$12 input/cached/output rates.
  • Benchmark-task spend is not your invoice. It excludes retries, review time, latency, subscriptions, and the cost of a wrong answer.
  • Deduced rows remain visible as dated research bets but cannot become recommendations.
Canonical 42-point table · data: August 13, 2026 · ordered by AA Intelligence Index, AA’s own summary of its whole board and not a routing input here · Execute is this card’s own score over four evaluations · deduced rows are bridge bets on record and cannot choose a route
Model · effortAA IndexExecute$ / execute task$ / knowledge taskEvidence
Claude Opus 5 · max63.0563.58$2.8238$0.2982printed
Claude Opus 5 · xhigh62.5263.04$2.1610$0.2252printed
Claude Fable 5 · max62.0761.23$3.6782$0.4465printed
Claude Opus 5 · high61.4862.10$1.4518$0.1536printed
Claude Fable 5 · xhigh61.34deduced
GPT-5.6 Sol · max60.9362.43$1.4925$0.2339printed
Grok 4.6 · high60.9263.76$1.1601$0.0590printed
Kimi K3 · max59.7062.20$0.9910$0.2408printed
Claude Fable 5 · high59.43deduced
GPT-5.6 Sol · xhigh59.0160.68$1.0189$0.1159printed
Claude Opus 5 · medium58.6457.87$0.8509$0.0821printed
Claude Fable 5 · medium57.84deduced
GPT-5.6 Sol · high57.3359.26$0.6957$0.0676printed
Claude Opus 4.8 · max57.3356.64$2.5607$0.4015printed
GPT-5.6 Terra · max56.5859.01$0.6159$0.1338printed
Grok 4.5 · high55.7657.27$0.5021$0.0539printed
GPT-5.6 Sol · medium55.5757.92$0.4871$0.0364printed
Claude Sonnet 5 · max55.2656.58$1.9641$0.4932printed
Claude Opus 4.8 · xhigh54.57deduced
Claude Fable 5 · low53.79deduced
DeepSeek V4 Pro 0813 · max53.0055.48$0.0642$0.0158printed
GPT-5.6 Terra · xhigh52.7753.76$0.3783$0.0499printed
Claude Opus 5 · low52.4650.63$0.5116$0.0281printed
GPT-5.6 Luna · max52.3254.65$0.0536$0.0181printed
Claude Opus 4.8 · high52.08deduced
DeepSeek V4 Flash 0731 · max51.7755.21$0.0333$0.0059printed
Gemini 3.6 Flash · high51.5851.55$0.7677$0.0466printed
Claude Sonnet 5 · xhigh51.03deduced
GPT-5.6 Sol · low50.7352.10$0.2991$0.0198printed
GPT-5.6 Terra · high50.1151.24$0.2738$0.0323printed
GPT-5.6 Luna · xhigh50.0652.00$0.0368$0.0099printed
Claude Opus 4.8 · medium50.06deduced
Claude Sonnet 5 · high48.88deduced
Kimi K3 · low48.2553.42$0.3560$0.0141printed
GPT-5.6 Luna · high46.9648.45$0.0262$0.0055printed
GPT-5.6 Terra · medium46.7648.20$0.1525$0.0139printed
Claude Opus 4.8 · low44.96deduced · extrapolated
Claude Sonnet 5 · medium42.58deduced · extrapolated
GPT-5.6 Terra · low41.3042.07$0.1226$0.0080printed
GPT-5.6 Luna · medium38.9138.88$0.0143$0.0018printed
GPT-5.6 Luna · low33.8533.66$0.0113$0.0009printed
Claude Sonnet 5 · low32.61deduced · extrapolated

Use the normal route when you want one answer quickly. Change the setup only when your available models or budget actually differ.