Jul 2026

Which LLM, at Which Effort? Building an Apples-to-Apples Answer

TL;DR
Vague specs, analysis, truth-seeking
Claude Opus 5 high. Medium on a budget, max for high stakes.
Inside Claude Code
Claude Opus 5 high.
Simple, high-volume work
GPT-5.6 Luna max. Never below high.
Defined, checkable work
GPT-5.6 Sol xhigh. Verify the artifacts.
One best answer
Claude Opus 5 max. Claude Fable 5 max if a wrong fact costs more.
Top of the printed board
60.7AA Index
Claude Opus 5 max · $2.03/task
Within 3 points, for less
49%cheaper
GPT-5.6 Sol max · 58.9 at $1.04
Cheapest measured config
$0.02per task
DeepSeek V4 Flash max · 40.3 Index
Spread, cheapest to top
91×on price
for 20.4 Index points of quality

Printed values only, July 24, 2026. Every figure here is computed from the same rows the charts plot.

Pick your model

Your modelsAll 12. Deselect what you can't use.

On a Codex-only stack, deselect the rest. Same for models your company blocks. The tree and charts recompute for what's left; the written explanations always describe the full board.

Four questions, in order. The first yes wins.

  1. 01

    Is the task fuzzy: specs from vague requirements, analysis, or deciding what's true?

    Claude Opus 5· high

    Claude Opus 5 is first on all three work-product boards: combined Briefcase, Analytical Quality, and GDPval. Every one of those leads clears the confidence intervals. Its rungs own the judgment frontier from medium up, and it is the only model here scored at every rung on both judgment components: step down to medium when budget binds, up to xhigh or max when stakes are high. Switch to Claude Fable 5 when knowledge is the binding constraint. Fable still leads Omniscience 40.15 to 31.27 and retakes the judgment lead at knowledge-heavy weights. On a tight budget, Grok 4.5 high scores 63.9 at about 6% of Opus 5 max's judgment bill.

  2. 02

    Are you inside Claude Code?

    Claude Opus 5· high

    GPT-5.6 Sol doesn't run there. Claude Opus 5 high is the vendor's own default effort; step up to xhigh for long agentic work. On measured Intelligence-Index values Opus 5 beats Claude Fable 5 outright at both max and xhigh, and costs less at each. Below that the comparison runs against Fable projections carrying a ±1.3 bar, so treat it as a point-estimate read rather than a settled one.

  3. 03

    Is it simple and high-volume, such as classification, extraction, or boilerplate?

    GPT-5.6 Luna· max

    Max prints 51.24 Index at $0.21 per Intelligence-Index task. Max by default; xhigh or high when a lower quality bar clears, never below high. The cliff below high is printed, not estimated.

  4. 04

    Everything else: a defined task you can verify?

    GPT-5.6 Sol· xhigh

    Xhigh prints 57.65 Index at $0.68 per Intelligence-Index task, the last escalator step before the price of a quality point doubles again. Verify completions by artifacts, not by the model's word, and drop to high for cheap bulk. GPT-5.6 Sol max still edges Claude Opus 5 high on both axes, but by 0.026 quality points and under 2% on cost. That is a razor margin, not a reason to switch mid-task.

Escape hatchWhen a single answer must be the best available, use the model at the top of the printed board, unless the answer turns on knowledge rather than analysis, where Claude Fable 5 max still leads.Claude Opus 5· max

The two frontiers

The first chart asks what quality costs. The second asks how good the judgment is, meaning analysis plus getting facts right, priced on its own workload of Briefcase tasks.

The execution frontier
data: July 24, 2026
30354045505560$0.01$0.02$0.05$0.1$0.2$0.5$1$2Cost per Intelligence-Index task (USD, log scale)AA Intelligence IndexClaude Opus 5 maxDeepSeek V4 Flash max
OpenAI
SolTerraLuna
Anthropic
Opus 5Fable 5Opus 4.8Sonnet 5
Moonshot
Kimi K3
DeepSeek
DeepSeek V4 ProDeepSeek V4 Flash
xAI
Grok 4.5
Google
3.6 Flash
colour = vendor, lighter = lower rung · filled = both axes from the evaluator · hollow = projected or estimated · bars = shown uncertainty
All 49 visible configurations. The dashed line marks the settings nothing else beats on both quality and price at once; it threads point estimates, so neighbors inside each other's shown bands are effectively tied. Hollow dots are projections or tier-4 estimates. Projected costs also carry roughly ±40% cross-model shape uncertainty, which is not drawn as a horizontal bar.Scope: standard reasoning-dial configurations only. For GPT-5.6 that means Artificial Analysis's low-to-max standard-mode rows, excluding none and the separate pro mode.
The judgment frontierdata: July 24, 2026

A different cost basis: the Briefcase judgment workload, so Claude Opus 5 max is $17.79/task here, not the $2.03 Intelligence-Index task cost above.

020406080100$0.05$0.1$0.5$1$5$10$20Cost per Briefcase task (USD, log scale)Spec Index (0–100)
OpenAI
Sol
Anthropic
Opus 5maxOpus 5xhighOpus 5highOpus 5mediumOpus 5lowFable 5Opus 4.8Sonnet 5
Moonshot
Kimi K3
DeepSeek
DeepSeek V4 ProDeepSeek V4 Flash
xAI
Grok 4.5
Google
3.6 Flash
on the frontier dominated: Claude Fable 5, Kimi K3, GPT-5.6 Sol, Claude Opus 4.8, Claude Opus 5 low, Claude Sonnet 5, Gemini 3.6 Flash frontier
Spec quality blends Briefcase analytical-quality Elo and the Omniscience knowledge index, each scaled across the scored set. DeepSeek V4 Flash is 0 because it is last on both inputs. No row reaches 100 because Claude Opus 5 tops analysis while Claude Fable 5 tops knowledge. Use the score for ordering and rungs. Claude Opus 5 owns the frontier from medium effort up, so you buy down its own ladder rather than switching models. Claude Fable 5 sits just off it and takes the lead once knowledge carries about 61% of the weight, which is the one substitution this chart still supports.

What they show:

  • Dropping Claude Opus 5 from max to low costs ten Index points and saves 82% of the bill. Its printed dial runs 60.69 Index at $2.03 per Intelligence-Index task down to 50.61 at $0.36, five printed rungs.
  • Claude Fable 5 max, the previous top, is beaten on both axes twice over. It prints 59.86 Index at $2.75. Claude Opus 5 max scores 0.83 Index points higher for 74% of that price; Opus 5 xhigh scores 0.21 higher for 57%. Better and cheaper at once is what dominated means.
  • DeepSeek holds the cheap floor. V4 Pro-high prints 43.11 Index at $0.0405 a task. Below roughly nine cents, no other vendor has a printed configuration that competes.
  • Claude Opus 5 at medium scores 75.9 for $5.25 per Briefcase task; max scores 93.0 for $17.79. Those judgment scores are my own composite, not a published benchmark: Briefcase analytical quality and the Omniscience knowledge index are each scaled across the scored set, then blended half and half, so the number orders the options rather than measuring them.
The uncomfortable one

Claude Opus 4.8's measured configuration is dominated on printed values alone, and now by its own successor.

Three Claude Opus 5 settings, three GPT-5.6 Sol settings, and Kimi K3 all deliver equal or better quality for less than Claude Opus 4.8's price. On the judgment frontier the same-vendor gap is starker: Opus 5 at medium effort scores 12.7 points higher than Opus 4.8 at max, for 64% of the bill. The printed-only claim needs no estimate. Toggle the chart and watch it hold.

That's the answer.

Everything below is evidence, method, and limits. Open what you want to check, skip the rest.

The best upgrade to this analysis isn't on any leaderboard and costs about two hours: ten tasks from your own backlog, run through the tree's top picks, judged by your own bar.

The receipts

Full data tableall 49 configurations, sorted by quality

Every included standard reasoning-dial configuration, sorted by quality: 30 printed, 12 projected, 0 cost-estimated, and 7 tier-4 estimated. For GPT-5.6, this means AA's standard-mode low-to-max rows, not none or pro. Claude Opus 5's dial is low/medium/high/xhigh/max, default high, and AA scored all five; its thinking-disabled mode ships only at low, medium, and high, and AA scores no row for it. Gemini 3.6 Flash's dial is minimal/low/medium/high; AA scored its top setting, high. Printed means both axes come from the current Artificial Analysis payload. Projected rows use vendor effort curves, with quality bars and roughly ±40% cost-shape uncertainty. Tier-4 means both axes come from cross-model dial behavior, the weakest grade, with bands on both axes.

Canonical 49-point table · data: July 24, 2026
Model · effortAA IndexCost / Intelligence-Index taskEvidence
Claude Opus 5 · max60.69$2.0277printed
Claude Opus 5 · xhigh60.07$1.5612printed
Claude Fable 5 · max59.86$2.7498printed
Claude Fable 5 · xhigh59.64 ±1.3~$1.8092projected
Claude Fable 5 · high58.90 ±1.3~$1.2009projected
GPT-5.6 Sol · max58.89$1.0373printed
Claude Opus 5 · high58.86$1.0571printed
Claude Fable 5 · medium58.44 ±1.3~$0.8323projected
GPT-5.6 Sol · xhigh57.65$0.6825printed
Kimi K3 · max57.11$0.9541printed
Claude Opus 5 · medium56.28$0.6184printed
Claude Fable 5 · low55.87 ±1.3~$0.5235projected
GPT-5.6 Sol · high55.87$0.4530printed
Claude Opus 4.8 · max55.69$1.7972printed
Claude Opus 4.8 · xhigh55.13 ±2.0~$1.1824projected
GPT-5.6 Terra · max54.95$0.8246printed
Grok 4.5 · high53.83$0.3123printed
Claude Opus 4.8 · high53.67 ±2.0~$0.7849projected
GPT-5.6 Sol · medium53.59$0.3140printed
Claude Sonnet 5 · max53.35$1.5254printed
Kimi K3 · high~53.11 (51.11–55.28)~$0.4338 ($0.3892–$0.4974)tier-4 estimated
Claude Opus 4.8 · medium52.64 ±2.0~$0.5440projected
GPT-5.6 Terra · xhigh51.60$0.4769printed
Claude Sonnet 5 · xhigh51.48 ±4.3~$1.0036projected
GPT-5.6 Luna · max51.24$0.2094printed
Claude Opus 5 · low50.61$0.3607printed
Claude Opus 4.8 · low50.07 ±2.0~$0.3421projected
Gemini 3.6 Flash · high50.07$0.5013printed
Claude Sonnet 5 · high50.05 ±4.3~$0.6662projected
Grok 4.5 · medium~49.82 (47.83–52.00)~$0.1420 ($0.1274–$0.1628)tier-4 estimated
GPT-5.6 Sol · low49.44$0.1975printed
GPT-5.6 Luna · xhigh49.07$0.1390printed
GPT-5.6 Terra · high48.95$0.3363printed
Gemini 3.6 Flash · medium~47.50 (45.84–49.04)~$0.2990 ($0.2614–$0.3444)tier-4 estimated
Claude Sonnet 5 · medium47.16 ±4.3~$0.4617projected
GPT-5.6 Luna · high46.06$0.0948printed
GPT-5.6 Terra · medium45.57$0.1752printed
DeepSeek V4 Pro · max44.27$0.0448printed
Kimi K3 · low~44.12 (39.14–47.67)~$0.1785 ($0.1697–$0.1843)tier-4 estimated
Gemini 3.6 Flash · low~43.35 (39.55–46.52)~$0.1646 ($0.1392–$0.1890)tier-4 estimated
DeepSeek V4 Pro · high43.11$0.0405printed
Grok 4.5 · low~40.83 (35.85–44.38)~$0.0584 ($0.0556–$0.0603)tier-4 estimated
GPT-5.6 Terra · low40.47$0.1542printed
Claude Sonnet 5 · low40.45 ±4.3~$0.2904projected
DeepSeek V4 Flash · max40.28$0.0223printed
GPT-5.6 Luna · medium38.05$0.0504printed
DeepSeek V4 Flash · high37.46$0.0411printed
Gemini 3.6 Flash · minimal~37.07 (32.10–40.62)~$0.0938 ($0.0892–$0.0968)tier-4 estimated
GPT-5.6 Luna · low33.26$0.0405printed
The frontier, model by modelwho owns which price band, and where the margins are thin

On the printed frontier, the fourteen-to-twenty-one-cent shelf is contested. GPT-5.6 Luna-xhigh, GPT-5.6 Sol-low, and GPT-5.6 Luna-max interleave there; ownership flips twice within seven cents. Above that, Grok 4.5-high narrowly beats GPT-5.6 Sol-medium on both axes and owns about $0.31 to $0.45. GPT-5.6 Sol-high then owns up to $0.62, where Claude Opus 5-medium takes a narrow slice, and Sol resumes through $1.56. From there Claude Opus 5 owns the rest: xhigh to $2.03, then max.

Two of these margins are thin enough to name. GPT-5.6 Sol-max beats Claude Opus 5-high by 0.026 quality points and under 2% on cost. Grok 4.5-high beats GPT-5.6 Sol-medium by 0.238 points and half a percent. Both are real on the raw floats, and both sit inside the evaluator's own display rounding, which prints the first pair as 59 and 59 at $1.04 and $1.06. AA has already reprinted one of these costs once. Treat either relation as one reprint away from flipping.

Most models sit inside the frontier. DeepSeek holds the floor; Grok enters the middle.

  • DeepSeek V4 Pro and Flash: both high/max dials print on both axes. Pro-high sits on the frontier at 43.11/$0.0405; Pro-max follows at 44.27/$0.0448. Flash-high prints 37.46/$0.0411, while its full Intelligence Index run cost was $53.94. Flash-max and Pro-high dominate it. Both measured costs fell outside my recorded estimate bands. The grading record keeps those failed estimates so later measurements can test the method.
  • Claude Opus 5: its dial is low/medium/high/xhigh/max, default high, and AA scored every rung. Max and xhigh both dominate Claude Fable 5's printed point; medium takes a narrow band of its own; high and low are dominated, high by GPT-5.6 Sol-max on a razor margin and low by Sol-medium, Luna-max, and Grok-high. Its thinking-disabled mode ships at low, medium, and high only, and Anthropic's docs say requests that disable thinking at xhigh or max return an error. AA scores no row for that mode, so it stays named and off the chart.
  • GPT-5.6 Terra: dominated at every setting. Claude Opus 5-medium and GPT-5.6 Sol beat it from above, while GPT-5.6 Luna, DeepSeek, and Grok 4.5 beat it from below.
  • Kimi K3: its vendor dial is low/high/max. Max is the one Index-scored reasoning configuration and uses the vendor default; low and high are tier-4 estimates with bands on both axes. Max lands just under GPT-5.6 Sol-xhigh at a higher price, so it is dominated on cost and quality. It ranks fifth of 56 on combined Briefcase and still point-leads Claude Fable 5 on analytical quality with overlapping intervals, though three Claude Opus 5 rungs now sit above both.
  • Grok 4.5: high is the top of its dial and its one Index-scored reasoning configuration. It knocks GPT-5.6 Sol-medium off the frontier by a quarter of a quality point and half a percent on cost, thinner than the evaluator's display rounding. A documented 0.4% AA cost reprint did not change that relation. On the judgment frontier it holds the cheap-triage rung outright, beating even Claude Opus 5 at low effort.
  • Claude Sonnet 5: its max setting, 53.4 quality at $1.53, is worse than GPT-5.6 Sol-medium at nearly five times the price. Nine printed configurations beat it on both axes. Its four lower settings are projections from HLE and CursorBench effort curves with a wide error bar.
  • Gemini 3.6 Flash: high is the top of its minimal/low/medium/high thinking dial and its one Index-scored configuration, printing 50.07/$0.50. Six printed configurations beat it on both axes, so it sits inside the frontier; the three lower settings are tier-4 estimates that fall deeper in. Its judgment profile splits: tenth on the Omniscience board, above Sol, but analytical quality at the floor of this comparison, above only DeepSeek V4 Flash, so the composite still ranks it low.

Effort pricing climbs as an escalator, but the shape is per-model. Along GPT-5.6 Sol's fully printed dial, each step roughly doubles the price of a marginal quality point, about three cents per point at the bottom and twenty-nine at the top, and the final step to max buys barely more than a single point. Claude Opus 5's printed dial runs the other way: $0.045, then $0.170, $0.419, and $0.748, so the steps multiply by 3.74, then 2.47, then 1.79. The expensive jump is out of low effort, and the ratio flattens toward the top. Two fully measured dials, two different shapes, so treat the doubling rule as Sol's, not a law. Either way there's no magic setting where the economics break; the question is which step your task's quality bar requires.

Claude Opus 5 takes all three work-product boards. Combined Briefcase Elo puts it at 1720.43 [1707.96–1733.73], first of 56. Analytical quality: 2015.89 [1991.24–2042.50], first. GDPval, which scores end-to-end work: 1860.57 [1835.37–1885.76], first of 180. Each of those leads clears the intervals of the model it displaced.

The old order survives intact underneath. Claude Fable 5 is fourth on combined Briefcase at 1573.78 [1562.40–1585.03] and third on GDPval at 1746.74 [1729.38–1764.11]; Kimi K3 1541.34 [1530.17–1551.50] and GPT-5.6 Sol 1504.82 [1494.03–1515.65] follow at fifth and sixth, both adjacent leads clearing the intervals. GDPval still puts Fable and Sol 10.99 points apart with overlapping intervals, and Sol is cheaper. On analytical quality Kimi's 1750.94 point-leads Fable's 1739.68 in a statistical tie, and Fable leads Sol's 1599.67 by 140.01. Sonnet sub-max has Briefcase rows but no Omniscience rows, so the composite scores only Sonnet-max. Fable's judgment rows remain max-only, while Opus 5 is measured at every effort on all three boards.

Knowledge is the one board Claude Fable 5 keeps. Omniscience ranks Fable first at 40.15, Gemini 3.1 Pro Preview second, Claude Opus 5 third at 31.27, Claude Opus 4.8 sixth, and Grok 4.5 seventh; Gemini 3.6 Flash is 10th, Sol 13th, Kimi 19th, Terra 68th, and Luna 107th. Fable's lead comes from knowing more, not from answering less. Its raw hallucination rate is 0.55, worse than Opus 4.8 at 0.36 and Kimi at 0.51; Opus 5 sits near 0.50 at every effort. Sol is the extreme at 0.89. METR also reports a high detected-cheating rate for Sol on one ReAct harness, though METR calls the numbers non-robust.

Judgment overrides the chart

The cost-quality chart alone cannot settle the fuzzy-task pick: Claude Opus 5 and Claude Fable 5 lead different boards.

Route by which failure costs you more. If bad analysis is the expensive one, take Claude Opus 5. If a confident wrong fact is, take Claude Fable 5.

How the numbers were builtone neutral harness, two axes, four evidence grades

One neutral harness, two separate axes. Artificial Analysis runs Intelligence Index v4.1 in uniform scaffolding across vendors. Quality is the Index score. Cost is its benchmark-weighted model bill per Intelligence-Index task run, based on token use and token prices. It is not conditioned on success and includes neither retries nor human time. Per-token price can still mislead because models consume different amounts of input, cached, reasoning, and answer tokens on the same task; per-task cost captures that token appetite, while the quality axis captures performance. Claude Fable 5's printed point is the evaluator's adaptive-reasoning configuration with automatic Claude Opus 4.8 fallback.

Why not just read each vendor's own benchmarks? Because the harness moves scores as much as the model does. The same Claude Opus 4.5 swings 9.5 points on SWE-bench Pro between a standardized harness and a vendor one. A cross-vendor gap measured under different harnesses is noise wearing a number, so only one neutral evaluator's runs compare.

Stop comparing LLM agents without disclosing the harness.

Source: arXiv preprint 2605.23950 (May 2026), source of the 9.5-point same-model swing

Every point sits at the best evidence grade it can honestly hold: printed on both axes; projected from vendor effort curves; evaluator-flagged quality with cost inferred from measured dials; or, weakest, a tier-4 cross-model estimate. Bands show measured donor spread or recipe sensitivity. They do not represent full confidence intervals. Tier names are never matched across vendors, and a setting a vendor doesn't ship is never invented. The printed only toggle strips everything below the top grade.

The judgment frontier is a composite. It blends Briefcase analytical-quality Elo with the Omniscience knowledge-reliability index. Each component is scaled across the compared set and blended half and half. Sweeping the blend weight from one quarter to three quarters forms the band; cost comes from Briefcase's printed workload bills. Use it for ordering and rungs. Every dominance verdict is checked against the weight sweep, Elo confidence intervals, and dropping any single model.

Model · effortSpec / 100Cost / Briefcase taskFrontier
Claude Opus 5 · max93.0 (89.496.5)$17.79on
Claude Opus 5 · xhigh91.5 (87.695.5)$14.26on
Claude Fable 5 · max89.0 (83.594.5)$22.30dominated
Claude Opus 5 · high86.0 (83.588.6)$10.41on
Claude Opus 5 · medium75.9 (75.476.5)$5.25on
Kimi K3 · max72.2 (68.975.6)$10.69dominated
GPT-5.6 Sol · max68.8 (67.869.8)$5.43dominated
Grok 4.5 · high63.9 (56.871.0)$1.12on
Claude Opus 4.8 · max63.2 (54.971.5)$8.26dominated
Claude Opus 5 · low60.5 (54.366.8)$1.78dominated
Claude Sonnet 5 · max58.2 (57.059.4)$14.43dominated
Gemini 3.6 Flash · high38.7 (21.256.2)$1.62dominated
DeepSeek V4 Pro · max15.8 (13.418.1)$0.08on
DeepSeek V4 Flash · max0.0$0.03on
Other models and limitswider comparison set and the limits that bound every claim

Other priced configurations. This comparison covers twelve chosen models. Artificial Analysis lists about ninety more priced configurations outside that scope. Examples include Gemini 3.5 Flash, GPT-5.5 at xhigh, and GLM-5.2 at max. Each is dominated by a charted printed configuration. The rest are not surveyed or judged here.

The evidence limits change how you read the dots. Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5 below max are projections. Kimi K3, Grok 4.5, and Gemini 3.6 Flash sub-max are weaker tier-4 estimates. No row uses the cost-estimated class. Per-task costs are benchmark-sized and scale with your own tasks. Values and evidence grades can move within weeks. The judgment frontier also blends two evaluator benchmarks into a score of my own design. Use it to order the options. The finer print:

  • Claude Opus 5's four sub-max rungs rest on a single capture. AA published them during the day I worked the model up: an archive snapshot of the same page two hours earlier had only max scored, xhigh and high present as bare directory entries, and no medium or low row at all. Max is verified across two capture times; the other four are corroborated across seven payloads fetched at the same moment, which is agreement between sources rather than over time. A model scored on its release day is exactly where an evaluator reprint is most likely, and the headline claims run through those rungs.
  • Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5 below max are projections, never neutral Index measurements. Fable-medium is 58.44 ±1.3 / ~$0.8323, with roughly ±40% cost-shape uncertainty; the judgment benchmarks score Fable-max only.
  • Claude Fable 5's printed numbers include its automatic fallback to Claude Opus 4.8 on a share of tasks; the evaluator plots the with-fallback configuration, and this post uses it consistently.
  • The Claude Opus 4.8 dominance claim is solid on printed values at its measured configuration and applies only at point estimates below it.
  • Projected costs borrow GPT-5.6 Sol's dial shape, so cross-model shape uncertainty remains. Claude Opus 5 supplies the first out-of-sample test of that borrowing: against its real printed cost dial, Sol's shape errs by 15%, 16%, 1%, and 7% at xhigh, high, medium, and low. That is inside the stated ±40%, but it tests the assumption on one model, not the projections themselves, which are for different models and stay ungraded.
  • Adding Claude Opus 5's dial to the tier-4 donor pool superseded every tier-4 estimate on this chart. Three of the seven midpoints moved outside the band the previous donor set published. That is why no tier-4 number should be read without its band.
  • The chart covers GPT-5.6 standard-mode low through max. It excludes none under AA's non-reasoning class and the orthogonal pro mode, for which AA has no rows. OpenAI's model-specific guide does not list legacy minimal; whether the API rejects or aliases it is untested.
  • Speed isn't on the chart and matters: GPT-5.6 Luna's printed token rate runs about three times GPT-5.6 Sol's.
  • Judgment coverage varies: Briefcase prints no rows for GPT-5.6 Luna or GPT-5.6 Terra; GDPval prints both. Sonnet sub-max has Briefcase rows but no Omniscience rows. Fable judgment evidence remains max-only, while Claude Opus 5 is scored at every effort on all three boards. Luna's knowledge index is negative across its whole dial. Kimi K3, Grok 4.5, Sonnet 5, and Gemini 3.6 Flash each have one Intelligence-Index-scored reasoning configuration.
  • AA values and evidence grades are mutable. Dated captures document Terra and Grok cost reprints, two DeepSeek estimates replaced by measurements, and a same-time disagreement between Sol's evaluation-board and model-page values.
  • The analytical-quality gap comes from one evaluator's Elo panel and rubric scores. METR's cheating figure comes from one public-model harness and is non-robust by METR's own description.
  • The judgment frontier is my own composite of two evaluator boards. Min-max scaling makes it set-relative; checks cover the weight sweep, Elo intervals, and set membership. Its cost comes from the judgment workload rather than the main chart.
  • Claude Fable 5 spent most of June suspended under a US export-control directive before global access returned in July. Availability can move as fast as scores.

Frequently Asked Questions