Which LLM, at Which Effort? Building an Apples-to-Apples Answer
- Vague specs, analysis, truth-seeking
- Claude Opus 5 high. Medium on a budget, max for high stakes.
- Inside Claude Code
- Claude Opus 5 high.
- Simple, high-volume work
- GPT-5.6 Luna max. Never below high.
- Defined, checkable work
- GPT-5.6 Sol xhigh. Verify the artifacts.
- One best answer
- Claude Opus 5 max. Claude Fable 5 max if a wrong fact costs more.
Printed values only, July 24, 2026. Every figure here is computed from the same rows the charts plot.
Pick your model
Your modelsAll 12. Deselect what you can't use.
On a Codex-only stack, deselect the rest. Same for models your company blocks. The tree and charts recompute for what's left; the written explanations always describe the full board.
Four questions, in order. The first yes wins.
- 01
Is the task fuzzy: specs from vague requirements, analysis, or deciding what's true?
Claude Opus 5· highClaude Opus 5 is first on all three work-product boards: combined Briefcase, Analytical Quality, and GDPval. Every one of those leads clears the confidence intervals. Its rungs own the judgment frontier from medium up, and it is the only model here scored at every rung on both judgment components: step down to medium when budget binds, up to xhigh or max when stakes are high. Switch to Claude Fable 5 when knowledge is the binding constraint. Fable still leads Omniscience 40.15 to 31.27 and retakes the judgment lead at knowledge-heavy weights. On a tight budget, Grok 4.5 high scores 63.9 at about 6% of Opus 5 max's judgment bill.
- 02
Are you inside Claude Code?
Claude Opus 5· highGPT-5.6 Sol doesn't run there. Claude Opus 5 high is the vendor's own default effort; step up to xhigh for long agentic work. On measured Intelligence-Index values Opus 5 beats Claude Fable 5 outright at both max and xhigh, and costs less at each. Below that the comparison runs against Fable projections carrying a ±1.3 bar, so treat it as a point-estimate read rather than a settled one.
- 03
Is it simple and high-volume, such as classification, extraction, or boilerplate?
GPT-5.6 Luna· maxMax prints 51.24 Index at $0.21 per Intelligence-Index task. Max by default; xhigh or high when a lower quality bar clears, never below high. The cliff below high is printed, not estimated.
- 04
Everything else: a defined task you can verify?
GPT-5.6 Sol· xhighXhigh prints 57.65 Index at $0.68 per Intelligence-Index task, the last escalator step before the price of a quality point doubles again. Verify completions by artifacts, not by the model's word, and drop to high for cheap bulk. GPT-5.6 Sol max still edges Claude Opus 5 high on both axes, but by 0.026 quality points and under 2% on cost. That is a razor margin, not a reason to switch mid-task.
The two frontiers
The first chart asks what quality costs. The second asks how good the judgment is, meaning analysis plus getting facts right, priced on its own workload of Briefcase tasks.
A different cost basis: the Briefcase judgment workload, so Claude Opus 5 max is $17.79/task here, not the $2.03 Intelligence-Index task cost above.
What they show:
- Dropping Claude Opus 5 from max to low costs ten Index points and saves 82% of the bill. Its printed dial runs 60.69 Index at $2.03 per Intelligence-Index task down to 50.61 at $0.36, five printed rungs.
- Claude Fable 5 max, the previous top, is beaten on both axes twice over. It prints 59.86 Index at $2.75. Claude Opus 5 max scores 0.83 Index points higher for 74% of that price; Opus 5 xhigh scores 0.21 higher for 57%. Better and cheaper at once is what dominated means.
- DeepSeek holds the cheap floor. V4 Pro-high prints 43.11 Index at $0.0405 a task. Below roughly nine cents, no other vendor has a printed configuration that competes.
- Claude Opus 5 at medium scores 75.9 for $5.25 per Briefcase task; max scores 93.0 for $17.79. Those judgment scores are my own composite, not a published benchmark: Briefcase analytical quality and the Omniscience knowledge index are each scaled across the scored set, then blended half and half, so the number orders the options rather than measuring them.
Claude Opus 4.8's measured configuration is dominated on printed values alone, and now by its own successor.
Three Claude Opus 5 settings, three GPT-5.6 Sol settings, and Kimi K3 all deliver equal or better quality for less than Claude Opus 4.8's price. On the judgment frontier the same-vendor gap is starker: Opus 5 at medium effort scores 12.7 points higher than Opus 4.8 at max, for 64% of the bill. The printed-only claim needs no estimate. Toggle the chart and watch it hold.
Everything below is evidence, method, and limits. Open what you want to check, skip the rest.
The best upgrade to this analysis isn't on any leaderboard and costs about two hours: ten tasks from your own backlog, run through the tree's top picks, judged by your own bar.
The receipts
Full data tableall 49 configurations, sorted by quality
Every included standard reasoning-dial configuration, sorted by quality: 30 printed, 12 projected, 0 cost-estimated, and 7 tier-4 estimated. For GPT-5.6, this means AA's standard-mode low-to-max rows, not none or pro. Claude Opus 5's dial is low/medium/high/xhigh/max, default high, and AA scored all five; its thinking-disabled mode ships only at low, medium, and high, and AA scores no row for it. Gemini 3.6 Flash's dial is minimal/low/medium/high; AA scored its top setting, high. Printed means both axes come from the current Artificial Analysis payload. Projected rows use vendor effort curves, with quality bars and roughly ±40% cost-shape uncertainty. Tier-4 means both axes come from cross-model dial behavior, the weakest grade, with bands on both axes.
| Model · effort | AA Index | Cost / Intelligence-Index task | Evidence |
|---|---|---|---|
| Claude Opus 5 · max | 60.69 | $2.0277 | printed |
| Claude Opus 5 · xhigh | 60.07 | $1.5612 | printed |
| Claude Fable 5 · max | 59.86 | $2.7498 | printed |
| Claude Fable 5 · xhigh | 59.64 ±1.3 | ~$1.8092 | projected |
| Claude Fable 5 · high | 58.90 ±1.3 | ~$1.2009 | projected |
| GPT-5.6 Sol · max | 58.89 | $1.0373 | printed |
| Claude Opus 5 · high | 58.86 | $1.0571 | printed |
| Claude Fable 5 · medium | 58.44 ±1.3 | ~$0.8323 | projected |
| GPT-5.6 Sol · xhigh | 57.65 | $0.6825 | printed |
| Kimi K3 · max | 57.11 | $0.9541 | printed |
| Claude Opus 5 · medium | 56.28 | $0.6184 | printed |
| Claude Fable 5 · low | 55.87 ±1.3 | ~$0.5235 | projected |
| GPT-5.6 Sol · high | 55.87 | $0.4530 | printed |
| Claude Opus 4.8 · max | 55.69 | $1.7972 | printed |
| Claude Opus 4.8 · xhigh | 55.13 ±2.0 | ~$1.1824 | projected |
| GPT-5.6 Terra · max | 54.95 | $0.8246 | printed |
| Grok 4.5 · high | 53.83 | $0.3123 | printed |
| Claude Opus 4.8 · high | 53.67 ±2.0 | ~$0.7849 | projected |
| GPT-5.6 Sol · medium | 53.59 | $0.3140 | printed |
| Claude Sonnet 5 · max | 53.35 | $1.5254 | printed |
| Kimi K3 · high | ~53.11 (51.11–55.28) | ~$0.4338 ($0.3892–$0.4974) | tier-4 estimated |
| Claude Opus 4.8 · medium | 52.64 ±2.0 | ~$0.5440 | projected |
| GPT-5.6 Terra · xhigh | 51.60 | $0.4769 | printed |
| Claude Sonnet 5 · xhigh | 51.48 ±4.3 | ~$1.0036 | projected |
| GPT-5.6 Luna · max | 51.24 | $0.2094 | printed |
| Claude Opus 5 · low | 50.61 | $0.3607 | printed |
| Claude Opus 4.8 · low | 50.07 ±2.0 | ~$0.3421 | projected |
| Gemini 3.6 Flash · high | 50.07 | $0.5013 | printed |
| Claude Sonnet 5 · high | 50.05 ±4.3 | ~$0.6662 | projected |
| Grok 4.5 · medium | ~49.82 (47.83–52.00) | ~$0.1420 ($0.1274–$0.1628) | tier-4 estimated |
| GPT-5.6 Sol · low | 49.44 | $0.1975 | printed |
| GPT-5.6 Luna · xhigh | 49.07 | $0.1390 | printed |
| GPT-5.6 Terra · high | 48.95 | $0.3363 | printed |
| Gemini 3.6 Flash · medium | ~47.50 (45.84–49.04) | ~$0.2990 ($0.2614–$0.3444) | tier-4 estimated |
| Claude Sonnet 5 · medium | 47.16 ±4.3 | ~$0.4617 | projected |
| GPT-5.6 Luna · high | 46.06 | $0.0948 | printed |
| GPT-5.6 Terra · medium | 45.57 | $0.1752 | printed |
| DeepSeek V4 Pro · max | 44.27 | $0.0448 | printed |
| Kimi K3 · low | ~44.12 (39.14–47.67) | ~$0.1785 ($0.1697–$0.1843) | tier-4 estimated |
| Gemini 3.6 Flash · low | ~43.35 (39.55–46.52) | ~$0.1646 ($0.1392–$0.1890) | tier-4 estimated |
| DeepSeek V4 Pro · high | 43.11 | $0.0405 | printed |
| Grok 4.5 · low | ~40.83 (35.85–44.38) | ~$0.0584 ($0.0556–$0.0603) | tier-4 estimated |
| GPT-5.6 Terra · low | 40.47 | $0.1542 | printed |
| Claude Sonnet 5 · low | 40.45 ±4.3 | ~$0.2904 | projected |
| DeepSeek V4 Flash · max | 40.28 | $0.0223 | printed |
| GPT-5.6 Luna · medium | 38.05 | $0.0504 | printed |
| DeepSeek V4 Flash · high | 37.46 | $0.0411 | printed |
| Gemini 3.6 Flash · minimal | ~37.07 (32.10–40.62) | ~$0.0938 ($0.0892–$0.0968) | tier-4 estimated |
| GPT-5.6 Luna · low | 33.26 | $0.0405 | printed |
The frontier, model by modelwho owns which price band, and where the margins are thin
On the printed frontier, the fourteen-to-twenty-one-cent shelf is contested. GPT-5.6 Luna-xhigh, GPT-5.6 Sol-low, and GPT-5.6 Luna-max interleave there; ownership flips twice within seven cents. Above that, Grok 4.5-high narrowly beats GPT-5.6 Sol-medium on both axes and owns about $0.31 to $0.45. GPT-5.6 Sol-high then owns up to $0.62, where Claude Opus 5-medium takes a narrow slice, and Sol resumes through $1.56. From there Claude Opus 5 owns the rest: xhigh to $2.03, then max.
Two of these margins are thin enough to name. GPT-5.6 Sol-max beats Claude Opus 5-high by 0.026 quality points and under 2% on cost. Grok 4.5-high beats GPT-5.6 Sol-medium by 0.238 points and half a percent. Both are real on the raw floats, and both sit inside the evaluator's own display rounding, which prints the first pair as 59 and 59 at $1.04 and $1.06. AA has already reprinted one of these costs once. Treat either relation as one reprint away from flipping.
Most models sit inside the frontier. DeepSeek holds the floor; Grok enters the middle.
- DeepSeek V4 Pro and Flash: both high/max dials print on both axes. Pro-high sits on the frontier at 43.11/$0.0405; Pro-max follows at 44.27/$0.0448. Flash-high prints 37.46/$0.0411, while its full Intelligence Index run cost was $53.94. Flash-max and Pro-high dominate it. Both measured costs fell outside my recorded estimate bands. The grading record keeps those failed estimates so later measurements can test the method.
- Claude Opus 5: its dial is low/medium/high/xhigh/max, default high, and AA scored every rung. Max and xhigh both dominate Claude Fable 5's printed point; medium takes a narrow band of its own; high and low are dominated, high by GPT-5.6 Sol-max on a razor margin and low by Sol-medium, Luna-max, and Grok-high. Its thinking-disabled mode ships at low, medium, and high only, and Anthropic's docs say requests that disable thinking at xhigh or max return an error. AA scores no row for that mode, so it stays named and off the chart.
- GPT-5.6 Terra: dominated at every setting. Claude Opus 5-medium and GPT-5.6 Sol beat it from above, while GPT-5.6 Luna, DeepSeek, and Grok 4.5 beat it from below.
- Kimi K3: its vendor dial is low/high/max. Max is the one Index-scored reasoning configuration and uses the vendor default; low and high are tier-4 estimates with bands on both axes. Max lands just under GPT-5.6 Sol-xhigh at a higher price, so it is dominated on cost and quality. It ranks fifth of 56 on combined Briefcase and still point-leads Claude Fable 5 on analytical quality with overlapping intervals, though three Claude Opus 5 rungs now sit above both.
- Grok 4.5: high is the top of its dial and its one Index-scored reasoning configuration. It knocks GPT-5.6 Sol-medium off the frontier by a quarter of a quality point and half a percent on cost, thinner than the evaluator's display rounding. A documented 0.4% AA cost reprint did not change that relation. On the judgment frontier it holds the cheap-triage rung outright, beating even Claude Opus 5 at low effort.
- Claude Sonnet 5: its max setting, 53.4 quality at $1.53, is worse than GPT-5.6 Sol-medium at nearly five times the price. Nine printed configurations beat it on both axes. Its four lower settings are projections from HLE and CursorBench effort curves with a wide error bar.
- Gemini 3.6 Flash: high is the top of its minimal/low/medium/high thinking dial and its one Index-scored configuration, printing 50.07/$0.50. Six printed configurations beat it on both axes, so it sits inside the frontier; the three lower settings are tier-4 estimates that fall deeper in. Its judgment profile splits: tenth on the Omniscience board, above Sol, but analytical quality at the floor of this comparison, above only DeepSeek V4 Flash, so the composite still ranks it low.
Effort pricing climbs as an escalator, but the shape is per-model. Along GPT-5.6 Sol's fully printed dial, each step roughly doubles the price of a marginal quality point, about three cents per point at the bottom and twenty-nine at the top, and the final step to max buys barely more than a single point. Claude Opus 5's printed dial runs the other way: $0.045, then $0.170, $0.419, and $0.748, so the steps multiply by 3.74, then 2.47, then 1.79. The expensive jump is out of low effort, and the ratio flattens toward the top. Two fully measured dials, two different shapes, so treat the doubling rule as Sol's, not a law. Either way there's no magic setting where the economics break; the question is which step your task's quality bar requires.
Claude Opus 5 takes all three work-product boards. Combined Briefcase Elo puts it at 1720.43 [1707.96–1733.73], first of 56. Analytical quality: 2015.89 [1991.24–2042.50], first. GDPval, which scores end-to-end work: 1860.57 [1835.37–1885.76], first of 180. Each of those leads clears the intervals of the model it displaced.
The old order survives intact underneath. Claude Fable 5 is fourth on combined Briefcase at 1573.78 [1562.40–1585.03] and third on GDPval at 1746.74 [1729.38–1764.11]; Kimi K3 1541.34 [1530.17–1551.50] and GPT-5.6 Sol 1504.82 [1494.03–1515.65] follow at fifth and sixth, both adjacent leads clearing the intervals. GDPval still puts Fable and Sol 10.99 points apart with overlapping intervals, and Sol is cheaper. On analytical quality Kimi's 1750.94 point-leads Fable's 1739.68 in a statistical tie, and Fable leads Sol's 1599.67 by 140.01. Sonnet sub-max has Briefcase rows but no Omniscience rows, so the composite scores only Sonnet-max. Fable's judgment rows remain max-only, while Opus 5 is measured at every effort on all three boards.
Knowledge is the one board Claude Fable 5 keeps. Omniscience ranks Fable first at 40.15, Gemini 3.1 Pro Preview second, Claude Opus 5 third at 31.27, Claude Opus 4.8 sixth, and Grok 4.5 seventh; Gemini 3.6 Flash is 10th, Sol 13th, Kimi 19th, Terra 68th, and Luna 107th. Fable's lead comes from knowing more, not from answering less. Its raw hallucination rate is 0.55, worse than Opus 4.8 at 0.36 and Kimi at 0.51; Opus 5 sits near 0.50 at every effort. Sol is the extreme at 0.89. METR also reports a high detected-cheating rate for Sol on one ReAct harness, though METR calls the numbers non-robust.
The cost-quality chart alone cannot settle the fuzzy-task pick: Claude Opus 5 and Claude Fable 5 lead different boards.
Route by which failure costs you more. If bad analysis is the expensive one, take Claude Opus 5. If a confident wrong fact is, take Claude Fable 5.
How the numbers were builtone neutral harness, two axes, four evidence grades
One neutral harness, two separate axes. Artificial Analysis runs Intelligence Index v4.1 in uniform scaffolding across vendors. Quality is the Index score. Cost is its benchmark-weighted model bill per Intelligence-Index task run, based on token use and token prices. It is not conditioned on success and includes neither retries nor human time. Per-token price can still mislead because models consume different amounts of input, cached, reasoning, and answer tokens on the same task; per-task cost captures that token appetite, while the quality axis captures performance. Claude Fable 5's printed point is the evaluator's adaptive-reasoning configuration with automatic Claude Opus 4.8 fallback.
Why not just read each vendor's own benchmarks? Because the harness moves scores as much as the model does. The same Claude Opus 4.5 swings 9.5 points on SWE-bench Pro between a standardized harness and a vendor one. A cross-vendor gap measured under different harnesses is noise wearing a number, so only one neutral evaluator's runs compare.
“Stop comparing LLM agents without disclosing the harness.”
Source: arXiv preprint 2605.23950 (May 2026), source of the 9.5-point same-model swing
Every point sits at the best evidence grade it can honestly hold: printed on both axes; projected from vendor effort curves; evaluator-flagged quality with cost inferred from measured dials; or, weakest, a tier-4 cross-model estimate. Bands show measured donor spread or recipe sensitivity. They do not represent full confidence intervals. Tier names are never matched across vendors, and a setting a vendor doesn't ship is never invented. The printed only toggle strips everything below the top grade.
The judgment frontier is a composite. It blends Briefcase analytical-quality Elo with the Omniscience knowledge-reliability index. Each component is scaled across the compared set and blended half and half. Sweeping the blend weight from one quarter to three quarters forms the band; cost comes from Briefcase's printed workload bills. Use it for ordering and rungs. Every dominance verdict is checked against the weight sweep, Elo confidence intervals, and dropping any single model.
| Model · effort | Spec / 100 | Cost / Briefcase task | Frontier |
|---|---|---|---|
| Claude Opus 5 · max | 93.0 (89.4–96.5) | $17.79 | on |
| Claude Opus 5 · xhigh | 91.5 (87.6–95.5) | $14.26 | on |
| Claude Fable 5 · max | 89.0 (83.5–94.5) | $22.30 | dominated |
| Claude Opus 5 · high | 86.0 (83.5–88.6) | $10.41 | on |
| Claude Opus 5 · medium | 75.9 (75.4–76.5) | $5.25 | on |
| Kimi K3 · max | 72.2 (68.9–75.6) | $10.69 | dominated |
| GPT-5.6 Sol · max | 68.8 (67.8–69.8) | $5.43 | dominated |
| Grok 4.5 · high | 63.9 (56.8–71.0) | $1.12 | on |
| Claude Opus 4.8 · max | 63.2 (54.9–71.5) | $8.26 | dominated |
| Claude Opus 5 · low | 60.5 (54.3–66.8) | $1.78 | dominated |
| Claude Sonnet 5 · max | 58.2 (57.0–59.4) | $14.43 | dominated |
| Gemini 3.6 Flash · high | 38.7 (21.2–56.2) | $1.62 | dominated |
| DeepSeek V4 Pro · max | 15.8 (13.4–18.1) | $0.08 | on |
| DeepSeek V4 Flash · max | 0.0 | $0.03 | on |
Other models and limitswider comparison set and the limits that bound every claim
Other priced configurations. This comparison covers twelve chosen models. Artificial Analysis lists about ninety more priced configurations outside that scope. Examples include Gemini 3.5 Flash, GPT-5.5 at xhigh, and GLM-5.2 at max. Each is dominated by a charted printed configuration. The rest are not surveyed or judged here.
The evidence limits change how you read the dots. Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5 below max are projections. Kimi K3, Grok 4.5, and Gemini 3.6 Flash sub-max are weaker tier-4 estimates. No row uses the cost-estimated class. Per-task costs are benchmark-sized and scale with your own tasks. Values and evidence grades can move within weeks. The judgment frontier also blends two evaluator benchmarks into a score of my own design. Use it to order the options. The finer print:
- Claude Opus 5's four sub-max rungs rest on a single capture. AA published them during the day I worked the model up: an archive snapshot of the same page two hours earlier had only max scored, xhigh and high present as bare directory entries, and no medium or low row at all. Max is verified across two capture times; the other four are corroborated across seven payloads fetched at the same moment, which is agreement between sources rather than over time. A model scored on its release day is exactly where an evaluator reprint is most likely, and the headline claims run through those rungs.
- Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5 below max are projections, never neutral Index measurements. Fable-medium is 58.44 ±1.3 / ~$0.8323, with roughly ±40% cost-shape uncertainty; the judgment benchmarks score Fable-max only.
- Claude Fable 5's printed numbers include its automatic fallback to Claude Opus 4.8 on a share of tasks; the evaluator plots the with-fallback configuration, and this post uses it consistently.
- The Claude Opus 4.8 dominance claim is solid on printed values at its measured configuration and applies only at point estimates below it.
- Projected costs borrow GPT-5.6 Sol's dial shape, so cross-model shape uncertainty remains. Claude Opus 5 supplies the first out-of-sample test of that borrowing: against its real printed cost dial, Sol's shape errs by 15%, 16%, 1%, and 7% at xhigh, high, medium, and low. That is inside the stated ±40%, but it tests the assumption on one model, not the projections themselves, which are for different models and stay ungraded.
- Adding Claude Opus 5's dial to the tier-4 donor pool superseded every tier-4 estimate on this chart. Three of the seven midpoints moved outside the band the previous donor set published. That is why no tier-4 number should be read without its band.
- The chart covers GPT-5.6 standard-mode low through max. It excludes
noneunder AA's non-reasoning class and the orthogonalpromode, for which AA has no rows. OpenAI's model-specific guide does not list legacyminimal; whether the API rejects or aliases it is untested. - Speed isn't on the chart and matters: GPT-5.6 Luna's printed token rate runs about three times GPT-5.6 Sol's.
- Judgment coverage varies: Briefcase prints no rows for GPT-5.6 Luna or GPT-5.6 Terra; GDPval prints both. Sonnet sub-max has Briefcase rows but no Omniscience rows. Fable judgment evidence remains max-only, while Claude Opus 5 is scored at every effort on all three boards. Luna's knowledge index is negative across its whole dial. Kimi K3, Grok 4.5, Sonnet 5, and Gemini 3.6 Flash each have one Intelligence-Index-scored reasoning configuration.
- AA values and evidence grades are mutable. Dated captures document Terra and Grok cost reprints, two DeepSeek estimates replaced by measurements, and a same-time disagreement between Sol's evaluation-board and model-page values.
- The analytical-quality gap comes from one evaluator's Elo panel and rubric scores. METR's cheating figure comes from one public-model harness and is non-robust by METR's own description.
- The judgment frontier is my own composite of two evaluator boards. Min-max scaling makes it set-relative; checks cover the weight sweep, Elo intervals, and set membership. Its cost comes from the judgment workload rather than the main chart.
- Claude Fable 5 spent most of June suspended under a US export-control directive before global access returned in July. Availability can move as fast as scores.