Which model, at which effort?

Depends what the whole job is worth to you. Drag the line.

Figure it out
Claude Fable 5.1 max$11.93 / task · 72.7 points · Top of the ladder
Execute it
Claude Fable 5.1 max$9.18 / task · 70.4 points · Top of the ladder

Figure it out. The highest measured score for mixed investigation. Every cheaper rung gives something up.

Execute it. The highest measured score for coding agent work. Every cheaper rung gives something up.

Every measured configurationquality against cost · open ↓
Higher is better, left is cheaper. The dashed line joins the rungs on the rail above; the ring is the current pick.

See the trade-off

Mixed investigation · higher is better, left is cheaper

Measured frontiers

Figure score against cost per task, 55 measured configurations, log cost axis. Higher is better; left is cheaper.Scored on Briefcase analytical quality, HLE, AA-LCR 1.1, AA-Omniscience index; the flags mark the recommended and selected configurations.
203040506070$0.01$0.10$1.00$10.00API dollars / profile task (USD, log scale)Figure scoreRECOMMENDED

Claude Fable 5.1 · maxSelected · 72.68 points · $11.93 / task

Chart key & sources
OpenAI
GPT-6 AstraLunaSolTerraGPT-5.5
Anthropic
Fable 5.1Opus 5Sonnet 5Fable 5Opus 4.8
Moonshot
Kimi K3
DeepSeek
DeepSeek V4 Pro 0813DeepSeek V4 Flash 0731
xAI
Grok 4.6Grok 4.5
Google
3.8 Flash3.7 Flash3.5 Flash
Zhipu
GLM-5.3-FlashGLM-5.3
Alibaba
Qwen3.8 27B
Meta
Muse Spark 1.3Muse Spark 1.2Muse Spark 1.1
Pale dashed line = measured Pareto frontier · coloured line = one model's effort rungs · the flags mark the recommended and selected configurations · data: September 6, 2026
Models you can use33 of 33 included · open ↓

Turn off a provider or model you cannot use. The rail recomputes from what remains.

Select a provider or model to include it. Open effort settings to exclude individual configurations.

5/5 efforts · edit
7/7 efforts · edit
7/7 efforts · edit
7/7 efforts · edit
4/4 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
5/5 efforts · edit
6/6 efforts · edit
5/5 efforts · edit
5/5 efforts · edit
5/5 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
2/2 efforts · edit
2/2 efforts · edit
4/4 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
2/2 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
How this is measuredsources, weights and limits · open ↓
How the rail is builtundominated rungs on a log job-share axis

The rail measures a whole job, not one benchmark task. Each dot is a model and effort that no cheaper measured configuration beats on the profile score. A dot sits at what that rung contributes to the job: its lane's share of the job, times what the rung costs against the best rung in its own lane. The right end of the rail is the job with both lanes at their best, so the ticks read 1, .1, .01 — fractions of that job, not dollars. Equal distances still mean equal cost ratios.

Why not one dollar figure for both lanes? Because a benchmark task is not a job, and the two lanes' tasks are not the same size. A Figure task is a set of hard questions and a business analysis, scored per question. An Execute task is a repository issue or a terminal session an agent works for minutes. Charging both lanes the same dollar ceiling treats one of each as an equal purchase — which it is not, when a real job runs one lane once and the other for hours.

"Your job is" sets the split. Mostly figuring puts 20% of the job in Execute, Balanced puts half, Mostly executing puts 80% there. The line is then one threshold both lanes' contributions sit under, and it converts back to a dollar ceiling per lane before anything is chosen — the lane carrying most of the job gets squeezed first, which is the point. Each answer still shows its own benchmark dollars per task and the date of the capture it came from.

Those three numbers are chosen, not measured. They are round positions, from the author's own working shape averaged informally with colleagues'. There is no retained source behind 20, 50 and 80, and none is claimed: they exist so a reader can say roughly what their week looks like, and a spinner would promise a precision nobody here has. Pick the one nearest your own and read the answer it gives.

The author's own jobs are a separate thing, and they are measured. Across 16 of his Claude Code jobs that ran both a figuring phase and an executing phase, 2026-09-03 to 09-08, executing took 5 to 94 percent of those two phases' API-equivalent cost at that day's published rates — median 84 percent, 75 percent weighted by dollars, over $347 of attributed spend. The largest such job spent $63 of $77 executing. That is one person's five days on one harness. It is not a bill, it is not a measurement of your work, and the width of that range is exactly why the split is a control rather than a constant.

Those runs were also almost all cache reads: 98.4 percent of input-side tokens across the Claude Code sessions, and 96.1 percent of input tokens across 258 Codex sessions in the same window. For work shaped like that a model's cache-read price orders the bill far more than its input price does, and a benchmark task bounded at a few minutes cannot show you that. Opus 5 reads cache at $0.50 per million tokens, Fable 5.1 at $0.25, Sonnet 5 at $0.20, Haiku 4.5 at $0.10 (Anthropic pricing, read 2026-09-08).

The rail carries the undominated rungs only. Cheaper and stronger steps are the dots themselves: every measured Pareto step for both lanes is on the line, and dragging to one selects it. A ceiling below every rung produces no recommendation for that lane and names the cheapest measured option instead.

How alternatives qualifyscreening rules and evidence limits

A model does not need every benchmark to earn consideration. One strong, relevant evaluation can qualify a scoped alternative. Figure looks across practical analysis, reasoning, factual reliability and document work. Execute looks across coding-agent work, professional outputs and tool workflows. A specialist remains a specialist; its best result does not establish broad capability.

Within your available models and effort settings, each benchmark admits results within 5 score points of its leader, or 100 Elo on native Elo boards. It also admits undominated quality–cost options within 20 points, or 400 Elo, of the leader. “Undominated” means no measured option is both at least as good and at least as cheap, with one strict advantage. Missing cost cannot earn a value claim. These margins are chosen screening rules, not uncertainty estimates or proof of a meaningful lead. Cheaper results below the quality floor remain in the dated discovery survey for follow-up.

The selected focus shows its core benchmarks and relevant supporting evidence. Coding agent work also shows FrontierCode, Zapier, AA's own Terminal-Bench run and SciCode; Deliverables shows Briefcase rubric checks and agentic professional work; Codebase investigation shows long-context reasoning and document reading. Supporting scores do not enter the profile average. Every qualifying model remains accessible; the first six are a compact alphabetical preview. Open a model’s row to see the exact efforts, benchmark scores, source, date, harness and fallback status. Several strong results from one publisher are not independent corroboration. Different effort settings on one card are separate observations, not a combined configuration.

The shortlist uses native Briefcase and GDPval Elo and the published AA-Omniscience index. Composite profiles use the fixed conversions below; the native evidence stays visible beside them. Native and composite units must not be compared directly. FrontierCode Fable runs have undisclosed fallback settings; Kimi's Zapier effort is undisclosed. These results can support scoped consideration, but cannot fill an exact broad-composite gap.

How the recommendation is pickedquality, spending and uncertainty

Start with quality, then decide how much to spend. Quality first selects the highest measured profile score among configurations with all required quality evidence, comparable API costs and reproducible settings. Equal scores prefer lower cost, then a stable configuration identity. It does not exchange score points for a guessed dollar value, and it does not claim to identify the best unmeasured model.

Choose Figure it out or Execute it, then narrow the work when a focused profile fits. A deliverables result can qualify on its own evidence even when the Coding Agent Index has not run that configuration and so cannot give it a broad Execute score. Undisclosed effort or fallback settings remain scoped observations; they cannot set the recommendation.

The rail sets a logarithmic ceiling in shares of one job: equal distances mean equal cost ratios, and the position converts to a dollar ceiling per lane through that lane's share of the job. Each kind of work then answers with its highest score within its own ceiling. Every measured Pareto step is a dot on the rail, so dragging past one selects it. A ceiling too low produces an explicit empty recommendation with the cheapest available alternative. Releasing the rail at its top removes the ceiling — Quality first — without changing model or effort exclusions.

Cheaper models that are not on the frontier still appear where their evidence earns it: the shortlist under Why this choice states each qualifying model’s measured scope, system, and unknown or above-ceiling costs. A component strength is not a broad recommendation.

The audit tests the policy, not the probability of being right. Open Method sensitivity & winning margins to compare half/double Elo slopes, wider group weights, omitted benchmarks/groups and focused compositions. Fixed cohorts keep the baseline configurations; expanded cohorts separately admit newly scoreable ones. Quality-only results allow unknown costs; priced results require comparable costs and apply your ceiling. Your access filters apply throughout. Scenario counts are not confidence percentages.

The reproducible all-model audit retains winners, coverage and native component margins. A stable result can still be wrong for your task: the scenarios cannot supply representative real-world outcomes that were never measured.

Why this choicemeasurements, method & access

55 of 83 evaluated configurations can set this profile’s recommendation.

Why 28 configurations cannot set this recommendation
  • Gemini 3.7 Flash low · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Gemini 3.7 Flash medium · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Gemini 3.8 Flash low · AA benchmark harness

    40.70 points · Comparable cost unavailable

    • Briefcase analytical quality: Direct task accounting is unavailable
  • Gemini 3.8 Flash medium · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Kimi K3 low · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Qwen3.8-Flash-Next xhigh · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Agnes 2.5 Pro Beta reasoning · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • GPT-5.5 low · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Qwen3.8 27B low · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Qwen3.8 27B medium · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • GPT-5.4 Pro xhigh · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • HLE: No published measurement at this exact configuration
    • HLE: Direct task accounting is unavailable
    • AA-LCR 1.1: No published measurement at this exact configuration
    • AA-LCR 1.1: Direct task accounting is unavailable
    • AA-Omniscience index: No published measurement at this exact configuration
    • AA-Omniscience index: Direct task accounting is unavailable
    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • GPT-5.5 Pro xhigh · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • HLE: No published measurement at this exact configuration
    • HLE: Direct task accounting is unavailable
    • AA-LCR 1.1: No published measurement at this exact configuration
    • AA-LCR 1.1: Direct task accounting is unavailable
    • AA-Omniscience index: No published measurement at this exact configuration
    • AA-Omniscience index: Direct task accounting is unavailable
    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Gemini 3.5 Flash minimal · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • DeepSeek V4 Flash Vision max · AA benchmark harness

    Score unavailable · Comparable cost unavailable

    • Briefcase analytical quality: No published measurement at this exact configuration
    • Briefcase analytical quality: Direct task accounting is unavailable
  • Claude Fable 5 high · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Fable 5 low · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Fable 5 medium · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Fable 5 xhigh · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Kimi K3 unspecified · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    Reasoning effort is not disclosed; cannot reproduce an exact setting

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • GPT-5.6 Luna none · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Opus 4.8 high · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Opus 4.8 low · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Opus 4.8 medium · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Opus 4.8 xhigh · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Claude Opus 5 none · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • Qwen3.8 Max unspecified · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    Reasoning effort is not disclosed; cannot reproduce an exact setting

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • GPT-5.6 Sol none · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
  • GPT-5.6 Terra none · Not published for this configuration

    Score unavailable · Comparable cost unavailable

    • HLE: Artificial Analysis has not published a model row for this configuration
    • HLE: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
    • Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
What the chosen configuration was measured on

The weights are declared policy, not a validated prediction of your work. Collaborative framing and planning remain unmeasured.

Score & native measurements

72.68 points · Mixed investigation. Benchmark proxies do not guarantee correctness.

BenchmarkNative resultAPI $ / task
HLE16.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)59.13%$1.59
AA-LCR 1.116.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)85.33%$1.58
AA-Omniscience index16.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)43.45 points$0.27
Briefcase analytical quality50.0% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)1965.99 Elo$22.71

Supporting evidence, outside the score: Briefcase presentation 1475.42 Elo · 2026-09-06. Same exact model and effort; AA benchmark harness.

Access: subscription, API or your hardware

Reviewed 2026-09-06. Availability does not reproduce a benchmark’s agent, effort or fallback configuration.

Subscription Verified route

Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start.

Plan-dependent inclusion or credits; the 50% is not extra capacity. Usage-based Enterprise uses API rates.

Claude Code requires 2.1.255 or newer. App fallback is automatic; availability alone does not reproduce the benchmark fallback system.

Fable plan rules · Fallback behavior

API Verified route

claude-fable-5-1

$10 input / $50 output USD per million tokens.

API fallback must be configured. Cost and quality depend on the exact fallback model/settings, not only the primary model. Effort: low through max; check native menu.

API pricing · Fable plan rules · Fallback behavior

Self-hosting Not verified

No verified open-weight or self-hosting route for this exact model.

Not verified

Do not infer exact-version availability or included usage from a family name.

API pricing · Fable plan rules · Fallback behavior

Method sensitivity & winning margins

These are declared scenarios, not confidence percentages. They vary Elo slope by half/double, group weights by 0.5/1/2, omit benchmarks/groups and test focused compositions. Your setup and exclusions apply. Priced results apply your spending ceiling; quality-only results ignore costs.

Fixed cohorts keep the baseline eligible configurations; expanded cohorts admit configurations newly scoreable under each composition. Neither fills missing scores. Several components share Artificial Analysis, model judges and agent tools; they are not independent confirmations. Benchmark-specific harness and fallback limits remain.

Baseline: 55 priced configurations from 24 models. Claude Fable 5.1 max leads Claude Fable 5.1 xhigh by 1.308 profile units.

  • HLE: native margin +0.0042 at 16.7% weight.
  • AA-LCR 1.1: native margin +0.0233 at 16.7% weight.
  • AA-Omniscience index: native margin +1.0667 at 16.7% weight.
  • Briefcase analytical quality: native margin +30.4100 at 50.0% weight.
ScenarioFixed · quality onlyFixed · pricedExpanded · quality onlyExpanded · priced
Declared profileClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Elo slope 0.5×Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Elo slope 2×Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×0.5, foundation ×0.5Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×0.5, foundation ×1Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×0.5, foundation ×2Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×1, foundation ×0.5Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×1, foundation ×1Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×1, foundation ×2Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×2, foundation ×0.5Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×2, foundation ×1Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Group weights analysis ×2, foundation ×2Claude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Without hleClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Without lcrClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Without omniscienceClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Without briefcaseAnalysisClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max67 configurations · 27 modelsClaude Fable 5.1 max67 configurations · 27 models
Without analysisClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max67 configurations · 27 modelsClaude Fable 5.1 max67 configurations · 27 models
Without foundationClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 modelsClaude Fable 5.1 max56 configurations · 24 modelsClaude Fable 5.1 max55 configurations · 24 models
Codebase investigation aloneNo comparable result0 configurations · 0 modelsNo comparable result0 configurations · 0 modelsClaude Fable 5.1 max47 configurations · 17 modelsClaude Fable 5.1 max47 configurations · 17 models

Evidence for mixed investigation

19 models

Relevant benchmark strengths, including setups that cannot set a priced recommendation. Supporting benchmarks stay separate from the profile score.

Agnes 2.5 Pro BetaLong-context reasoningMeasured with limits · One publisher

These strengths do not establish broad capability.

  • AA-LCR 1.1reasoning effort · 83.00 points · $0.012 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Claude Fable 5Scientific reasoning · Hard reasoning +1Comparable profile option · One publisher

At least one qualifying effort has complete, priced profile evidence.

Scientific reasoning · Hard reasoning · Factual reliability

  • CritPtmax effort · 28.57 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • HLEmax effort · 55.47 points · $0.93 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-Omniscience indexmax effort · 43.30 points · $0.070 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
Claude Fable 5.1Practical analysis · Scientific reasoning +3Comparable profile option · One publisher

At least one qualifying effort has complete, priced profile evidence.

Practical analysis · Scientific reasoning · Hard reasoning · Long-context reasoning · Factual reliability

  • Briefcase analytical qualitymax effort · 1965.99 Elo · $22.71 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • Briefcase analytical qualityxhigh effort · 1935.58 Elo · $17.79 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • Briefcase analytical qualityhigh effort · 1847.00 Elo · $11.40 / benchmark taskQuality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • CritPtmax effort · 29.71 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • CritPtxhigh effort · 31.14 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • CritPthigh effort · 30.29 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • CritPtmedium effort · 29.14 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • CritPtlow effort · 27.71 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • HLEmax effort · 59.13 points · $1.59 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • HLExhigh effort · 58.71 points · $1.03 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • HLEhigh effort · 55.93 points · $0.40 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • HLEmedium effort · 53.80 points · $0.23 / benchmark taskQuality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-LCR 1.1max effort · 85.33 points · $1.58 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-LCR 1.1high effort · 83.67 points · $1.48 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-LCR 1.1medium effort · 84.67 points · $1.48 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-Omniscience indexmax effort · 43.45 points · $0.27 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-Omniscience indexxhigh effort · 42.38 points · $0.077 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
  • AA-Omniscience indexhigh effort · 40.80 points · $0.021 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
Claude Opus 5Practical analysis · Scientific reasoning +2Comparable profile option · One publisher

At least one qualifying effort has complete, priced profile evidence.

Practical analysis · Scientific reasoning · Hard reasoning · Factual reliability

  • Briefcase analytical qualitymax effort · 1909.74 Elo · $17.79 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
  • Briefcase analytical qualityxhigh effort · 1900.45 Elo · $14.26 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
  • Briefcase analytical qualityhigh effort · 1815.72 Elo · $10.41 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
  • CritPtmax effort · 29.14 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
  • CritPtxhigh effort · 27.71 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
  • CritPthigh effort · 28.29 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
  • HLEmax effort · 54.87 points · $0.67 / benchmark taskNear the available leader · AA benchmark harness · single model · 2026-09-06
  • HLExhigh effort · 54.40 points · $0.52 / benchmark taskNear the available leader · AA benchmark harness · single model · 2026-09-06
  • AA-Omniscience indexmedium effort · 31.02 points · $0.0062 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Gemini 3.7 FlashFinancial analysis · Hard reasoningComparable profile option · One publisher

At least one qualifying effort has complete, priced profile evidence.

  • AA-AnalystAgent pass@1high effort · 70.50 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
  • HLEhigh effort · 47.87 points · $0.037 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Gemini 3.8 FlashHard reasoning · Long-context reasoning +1Measured with limits · One publisher

These strengths do not establish broad capability.

Hard reasoning · Long-context reasoning · Factual reliability

  • HLEmedium effort · 42.12 points · $0.025 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
  • AA-LCR 1.1medium effort · 84.00 points · $0.090 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
  • AA-Omniscience indexmedium effort · 28.58 points · $0.0036 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
How options qualify & what this search misses

Within your available configurations, each benchmark admits results within 5 points of its leader (100 for Elo), plus quality–cost frontier results within 20 points (400 Elo). These are screening choices, not confidence intervals. No cost means no value claim. Known weak results remain available in each source; a strength is not an endorsement of the whole model.

The 2026-09-06 survey examined 643 AA configurations, 95 FrontierCode entries and 95 Zapier entries. Coverage and public availability remain incomplete; this is a source-backed shortlist, not proof that every great model was found.

Access reference for all 33 maintained models
Access: subscription, API or your hardware

Reviewed 2026-09-06. Availability does not reproduce a benchmark’s agent, effort or fallback configuration.

Subscription Not verified

Vendor advertises a Token Plan; exact inclusion and pricing were not verified.

Not verified

Do not infer exact-version availability or included usage from a family name.

Vendor announcement · Vendor site

API Verified route

First-party Pro Beta API is announced by the vendor.

$0.10 input / $0.30 output per million tokens; cached input $0.01.

Pricing evidence is a vendor social post, not a verified API tariff page. Beta status remains; configurable reasoning efforts unverified.

Vendor announcement · Vendor site

Self-hosting Not verified

No verified open-weight or self-hosting route for this exact model.

Not verified

Do not infer exact-version availability or included usage from a family name.

Vendor announcement · Vendor site

Use your allowance wellchat, agent and when to switch

Practical guidance for each kind of work, from the configuration the rail names. There is no measured quota optimum.

Use your allowance well.

  1. Settle the decision

    Decide what success means with the model that has the useful context.

  2. Your agentClaude Fable 5.1 · max

    Start here when the investigation needs files, tests or tools.

Limits & when to switch

On Claude Pro and Max, Chat and Claude Code share usage limits. Switching between them adds no separate allowance.

Skip a separate chat when the plan is settled or the decision needs repository tools. Reconsider the model or effort when the design is unresolved or fixes keep failing; carry the useful context with you.

Use Cheaper to inspect the measured trade-off. API savings do not measure subscription savings, and the benchmarks do not identify when to escalate.

Selected model access: Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start. Check access & evidence below.

Product facts checked 2026-09-07. Claude Pro and Max usage

Practical guidance · no measured quota optimum.

Use your allowance well.

  1. Settle the decision

    Decide what success means with the model that has the useful context.

  2. Claude CodeClaude Fable 5.1 · max

    Bring the decision. Build, verify and fix in the same task.

Limits & when to switch

On Claude Pro and Max, Chat and Claude Code share usage limits. Switching between them adds no separate allowance.

Skip a separate chat when the plan is settled or the decision needs repository tools. Reconsider the model or effort when the design is unresolved or fixes keep failing; carry the useful context with you.

Use Cheaper to inspect the measured trade-off. API savings do not measure subscription savings, and the benchmarks do not identify when to escalate.

Selected model access: Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start. Check access & evidence below.

Product facts checked 2026-09-07. Claude Pro and Max usage

Practical guidance · no measured quota optimum.

Subscription, API or self-hostingwhat the cost numbers mean

Benchmark API cost is a comparison unit, not your bill. Subscription plans may include an allowance, require credits from the first request, restrict supported tools or leave exact versions undisclosed. The selected model’s access disclosure and the full access reference keep those distinctions separate, with official sources and verification dates.

Downloadable weights establish a self-hosting route, not laptop suitability or savings. The reference gives sourced serving examples and distinguishes author formats, community quantizations and related hosted variants. For example, hosted Qwen Flash is based on Flash-Next with additional production features; it is not silently treated as the same measured weights configuration. Local hardware, electricity and operations costs are not estimated here.

What turning the effort dial up costsmeasured steps and regressions

Turning the dial up buys the model more thinking. You pay for that thinking in tokens, so it is a spending dial first and a quality dial second.

That makes it measurable. Every row here is one model at one setting, scored on DeepSWE, Terminal-Bench 2.1 and SWE-Atlas-QnA and billed with the same job weights. The ladder compares complete measured settings; a missing setting is skipped, not estimated.

GPT-5.6 Sol shows it most starkly. Among models with at least four complete settings, the last step costs about 44× what the first did: $0.017 a point at the bottom, $0.74 at the top. Measured September 6, 2026.

GPT-5.6 Sol — each step of the reasoning dial, on the execution lane
StepScorePoints boughtAdded cost$ / point
nonelow55.23+11.86$0.2056$0.0173
lowmedium61.62+6.39$0.9005$0.1409
mediumhigh64.12+2.50$0.8075$0.3230
highxhigh63.34−0.78$0.7362bought nothing
xhighmax65.05+1.71$1.2588$0.7361

No number here tells you where to stop. The board prices a point at each step. What a wrong answer costs you is the other half, and only you have it.

Higher does not always mean better, either. 2 measured configurations score lower at a higher setting:

  • Claude Opus 5 · xhighmax 1.12 points down, $0.77 more per task. Paying more for a worse result.
  • GPT-5.6 Sol · highxhigh 0.78 points down, $0.74 more per task. Paying more for a worse result.

One board, one capture date, no error bars. Read that as a reason to check your own setting, not as a law about the model. Either way, “turn it to max” is a guess rather than advice.

Never pick the setting on its own. A cheap model turned up and a costly one turned down are the same kind of thing here — which is why the card compares configurations rather than ranking model names.

What each task is measured onsources, weights and missing evidence

A profile defines the work being compared. Four options, two per lane, each scored by one named source — so a focus can never quietly mix in evidence its own definition does not name. There used to be eight, which asked you which benchmark you trusted rather than what you were about to do.

LaneProfileRequired evidence and weights
Figure it outMixed investigationHalf reasoning/knowledge foundation, half Briefcase analytical quality — the AA model row, in points
Figure it outCodebase investigationSWE-Atlas-QnA alone, from the Coding Agent Index row for your harness, in percent correct
Execute itCoding agent workThe AA Coding Agent Index for your harness: DeepSWE, Terminal-Bench 2.1 and SWE-Atlas-QnA in equal thirds, in points
Execute itDeliverablesGDPval-AA v2 alone, in native Elo

What scores the Execute lane

The Execute lane reads the Artificial Analysis Coding Agent Index, captured at v1.4 on 2026-09-08, whose own page states it "incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA" in equal thirds, over three attempts per task.

A row on that board is a harness, a model and a reasoning setting together — "Claude Code / Fable 5.1 / max", not "Fable 5.1". That is the whole reason the lane moved to it: the number describes the pair you actually run, not a model in a scaffold nobody uses. Choosing Claude Code reads Claude Code's rows; choosing Codex reads Codex's; All models takes each configuration's best published row whatever agent produced it, and the answer names that agent. Cost is that row's own mean API dollars per task, cached input and cache writes included.

Its three parts: DeepSWE resolves repository issues end to end (113 tasks); Terminal-Bench v2.1 runs command-line coding, administration and data tasks (89 tasks); SWE-Atlas-QnA answers expert-written comprehension questions inside a repository (124 tasks). The last of those, alone and in percent, is what Codebase investigation scores on the Figure lane.

HLE tests hard academic questions; AA-LCR tests reasoning across document sets. Omniscience supplies the publisher's factual-reliability score. Briefcase separates analytical quality, rubric checks and presentation; its presentation score is never substituted for analysis. GDPval-AA v2 compares professional deliverables and is what Deliverables scores, alone, in its native Elo. FrontierCode Main and Zapier AutomationBench are now supporting evidence on the rows that carry them, not requirements: neither gates a lane any more.

Fixed conversions preserve differences without depending on the current roster. Composite profiles use:

Rate p                     → 100 × p
Published Omniscience O     → (O + 100) / 2
Native Briefcase/GDPval E   → (E − 500) / 20

Elo is not clipped: 1000 becomes 25 points, 1500 becomes 50 and 2500 becomes 100; higher or negative values remain possible. Twenty Elo per point is a declared cross-benchmark policy, not an empirically established exchange rate. 1000 Elo is not asserted to mean human-level performance. Single-benchmark profiles keep their native units. Adding or filtering models cannot move another configuration's score.

Costs use the same profile weights as quality. AA bills use per-benchmark input/output tokens, cache rates, prices and fallback fractions, divided by task count. Both Briefcase dimensions use its full task bill. The Coding Agent Index publishes its own mean API dollars per task per row, already inclusive of cached input and cache writes. FrontierCode publishes mean API dollars per rollout, already per task. Zapier publishes API dollars per task rounded to cents. Dedicated-deployment prices and incomplete fallback bills remain unpriced. Subscription inclusion is never substituted with zero benchmark cost.

Each profile requires all its components at the exact model, effort, version, harness and compatible system. Missing values stay unscored or unpriced; weights are never renormalized around them. A Coding Agent Index row is used only for the harness that produced it — a Codex number never wears a Claude Code label. A fallback-enabled AA system cannot borrow a standalone external result, even when observed fallback use is zero; the index's own fallback-enabled rows stay eligible and say so, the same rule the Figure lane applies to AA's. Focused source joins preserve model, effort, harness, fallback identity and source record. An estimated overall AA Index neither fills a gap nor excludes complete evidence.

What this composition costs you

Coverage. The index publishes 68 rows, of which 53 name a model compared here; they cover 49 model-and-effort configurations as of 2026-09-08 — 17 reachable under Claude Code and 21 under Codex once model compatibility is applied. A configuration the board has not run carries no Execute score, with that reason attached, and stays visible in the shortlist. Claude Fable 5.1 is published only at max, in Claude Code; GPT-6 Astra only at max, in Codex — so a reader running either below max is answered from the shortlist instead. The composition this one replaced required FrontierCode, Briefcase rubric, GDPval and Zapier all present, which admitted 29 configurations and excluded every Fable row — the reader's own model missing from the lane that decides their biggest bill.

The demoted sources still agree with it. Over the 31 configurations both measure, FrontierCode ranks them at Spearman 0.90 against the index; Zapier, over 28, at 0.89. Requiring FrontierCode was buying agreement the index already had, at the price of every configuration FrontierCode does not publish. Both remain on the rows.

Where the ordering comes from. Inside the index, Terminal-Bench 2.1 sits near its ceiling for the leading rows: across the ten highest-scoring rows in the 2026-09-08 capture it averages 86.6 against a board maximum of 91.0, while DeepSWE averages 65.0 and SWE-Atlas-QnA 49.2. The top of the board is therefore separated mostly by DeepSWE and SWE-Atlas-QnA, not by the terminal component.

The board moves. Between May and August 2026 the index went through four versions. Its benchmark line-up changed once, at v1.1 in June 2026, which added DeepSWE and removed SWE-Bench-Pro-Hard-AA. v1.2 and v1.3 changed how SWE-Atlas-QnA is scored — rubric reward, then binary pass/fail, then Scale's own Task Resolve Rate — and v1.4 upgraded Terminal-Bench from v2 to v2.1 and added reward-hacking detection (AA methodology, read 2026-09-08). Scores are not comparable across those versions. The extractor asserts the published version sentence, the exact set of dataset slugs and the equal-third weights, and stops the refresh rather than re-weighting the lane underneath you if any of them changes.

Its two coding parts each carry a caveat. DeepSWE (Datacurve, July 2026) is 113 original tasks over 91 repositories, graded binary with no partial credit, run for every model in one fixed harness — a deliberate trade its own paper describes as removing the scaffolding confound at the cost of cross-harness realism. It is also a single-vendor benchmark whose contamination controls cannot be audited from outside. SWE-Atlas-QnA (Scale AI, March 2026) is 124 questions over 11 production repositories; Scale's own leaderboard notes some configurations show an elevated refusal rate because the tasks trip security filters, which reads as a low score without being one.

Harness effects are the point, and also a limit. The number describes one agent at one version, and it does not transfer to the same model in another scaffold. That is not a small correction: a controlled 3×3 study on a 100-task SWE-bench Verified subset measured harness variance at 7.8× model variance and found rank reversals in 6 of 9 model-pair comparisons (arXiv 2605.23950, a position paper whose authors explicitly decline to claim that ratio is universal), and Harness-Bench found a 23.8-point aggregate gap between the best and worst configurable harnesses over the same 106 tasks and model pool (arXiv 2605.27922). Both were read 2026-09-08. This is the argument for scoring the pair you run — and the reason a number from one agent tells you little about another.

What the Figure lane's own sources are worth

The Execute lane is not the only one holding evidence with a record. Stated plainly, and dated:

  • HLE has contested answers. FutureHouse examined 321 text-only chemistry and biology questions and found that 29% ± 3.7 had answers contradicted by peer-reviewed literature (FutureHouse, 23 July 2025). The HLE team's own follow-up on that page put roughly 18% of a reviewed subset as problematic under three-expert review and committed to rolling revisions. The finding covers that subset, not the whole exam.
  • AA-Omniscience's questions are machine-written. Artificial Analysis states its 6,000 questions were generated by an LLM-based question-generation agent from authoritative academic and industry sources (AA, read 2026-09-08). The sources are human; the questions are not.
  • AA-LCR v1.1 corrected its own answers. AA's methodology records that v1.1 added a grading system prompt, corrected 16 answer keys and changed the grader, and that scores are not directly comparable with v1.0 (AA methodology, read 2026-09-08).
  • Briefcase is Artificial Analysis's own private benchmark. Its 91 tasks across four held-out scenarios were written by outside experts from firms including Google, McKinsey and BCG, but AA owns and runs it, only a fifth demonstration scenario is public, and AA's methodology table lists one repeat per task — no averaging over attempts (AA-Briefcase, read 2026-09-08). Its analytical-quality Elo is a preference ranking, which is not a correctness measure.

These are benchmark proxies, not success probabilities. Collaborative problem framing and planning remain unmeasured in Figure; selective escalation and supervision remain unmeasured in Execute. Agent harnesses, tools, task sets and judges differ. Shared AA provenance is not independent corroboration, and small numerical gaps need not matter in real work. A one-benchmark profile carries that scope next to its recommendation.

Data and limitscoverage, effort names, raw rows and sources

Which models are included

The comparison uses two dated captures: the Artificial Analysis model rows of September 6, 2026, which score the Figure lane and the Deliverables focus, and the Coding Agent Index board of September 8, 2026, which scores the Execute lane and Codebase investigation. The broad score table contains exact configurations with explicit lane coverage; the shortlist also includes individual benchmark observations that cannot form a full composite. Each result identifies its measured effort, harness and system; results do not transfer between efforts, harnesses or orthogonal configurations.

An older model stays while something here depends on it: a frontier step, a route, a qualifying shortlist strength or value option, a substitute for a named gap, or an ungraded prediction. Replacement within its vendor line starts that check; age alone never removes it.

Astra without reasoning remains excluded because current OpenAI documentation does not support that configuration. Muse Spark 1.3 max is now publicly available and included; missing costs remain visible. Some Muse native API effort mappings remain unverified, and DeepSeek Flash Vision is experimental. The shortlist states those limits beside each model. The discovery record retains availability checks, omissions and source evidence.

Supporting benchmarks and missing results

Additional evidence can qualify a scoped shortlist alternative without changing the broad score. Figure includes GDP.pdf, CritPt, AnalystAgent and ITBench. Execute includes FrontierCode, Zapier, AA's own Terminal-Bench 2.1 run, SciCode, EnterpriseOps-Gym, AA's partial-credit AutomationBench and APEX-Agents. AA's Terminal-Bench run and the Coding Agent Index's are the same tasks in different harnesses and are kept as separate measurements, never merged. Coverage varies across exact configurations; not measured is a missing result, not poor performance.

AA-Briefcase separately reports analytical quality and presentation. A low presentation score does not arithmetically lower its separate analytical-quality Elo. Its rubric, analytical Elo, and overall Elo answer different questions; none is silently substituted for another here.

Where the numbers come from

Artificial Analysis, its Coding Agent Index, Cognition, and Zapier supply the measurements and accounting. The retained source captures and reproduction commands show how those inputs become the two lane scores, direct costs, and recommendations; the Coding Agent Index capture retains the page as served, its hashes and what it does and does not support. Every page is parsed as data — no downloaded script is executed. Unknown source versions stop regeneration for review. Missing measurements or task accounting remain visible and prevent selection only in the affected lane.

Historical calibrated predictions remain in the dated research and falsifiability records. They do not enter the current comparison or choose a recommendation.

Effort dials, model by model

  • GPT-6 Astra has measured low, medium, high, xhigh, and max configurations here. OpenAI's model guide excludes none; Pro is a separate class.
  • Claude's effort documentation distinguishes model-specific effort settings. The measured Fable configurations include the source's fallback behavior; the direct bill accounts for those fractions.
  • Gemini Flash uses low, medium, and high thinking levels here. Each is independently measured.
  • DeepSeek and GLM document low / high / max; an unmeasured setting stays a named gap.
  • Qwen3.8-Flash-Next distinguishes its effort dial from disabling thinking. Only the measured xhigh configuration is placed here.
  • Agnes 2.5 Pro Beta uses reasoning here to name AA's measured class. That label is not an API effort value.

What the prices do not include

Benchmark API cost is not your invoice. It includes the evaluated runs' recorded API usage, but does not price your extra retries, human review, latency, subscriptions, or the cost of a wrong answer. An API cost ratio is not a subscription-quota ratio. The rail prices a benchmark task mix as shares of a job; it does not forecast the cost of your next job, and the job split it uses is a control you set, not a measurement of you.

83 catalogued configurations · September 6, 2026 · ordered by Figure score. Each lane uses its declared job weights. Missing evidence stays unscored or unpriced. These scores are benchmark proxies, not probabilities of success.
Model · effortFigure$ / figure taskExecute$ / execute task
Claude Fable 5.1 · max72.68$11.929970.43$9.1832
Claude Fable 5.1 · xhigh
Missing evidence
    • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
    • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
    • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
    71.37$9.3321UnscoredUnpriced
    Claude Opus 5 · max69.03$9.140167.03$8.9435
    Claude Opus 5 · xhigh68.75$7.344768.15$8.1708
    Claude Fable 5.1 · high
    Missing evidence
      • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
      • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
      • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
      68.67$6.0158UnscoredUnpriced
      Claude Fable 5.1 · medium
      Missing evidence
        • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
        • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
        • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
        66.72$4.5761UnscoredUnpriced
        Claude Opus 5 · high66.01$5.389165.62$3.9186
        GPT-6 Astra · max65.97$4.996466.97$4.7173
        GPT-6 Astra · xhigh
        Missing evidence
          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
          64.67$3.4940UnscoredUnpriced
          Claude Fable 5 · max64.32$11.580067.17$11.6874
          Claude Fable 5.1 · low
          Missing evidence
            • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
            • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
            • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
            63.66$3.6276UnscoredUnpriced
            GPT-6 Astra · high
            Missing evidence
              • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
              • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
              • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
              63.59$2.5732UnscoredUnpriced
              Muse Spark 1.3 · max62.93$1.924567.97$1.5842
              GPT-6 Astra · medium
              Missing evidence
                • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                62.55$2.1481UnscoredUnpriced
                GLM-5.3 · max
                Missing evidence
                  • Execute · DeepSWE: The Coding Agent Index does not publish this model
                  • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                  • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                  62.21$2.8899UnscoredUnpriced
                  Kimi K3 · max
                  Missing evidence
                    • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                    • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                    • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                    61.92$3.5065UnscoredUnpriced
                    Claude Opus 5 · medium61.89$2.779864.07$3.1719
                    Grok 4.6 · xhigh
                    Missing evidence
                      • Execute · DeepSWE: The Coding Agent Index does not publish this model
                      • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                      • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                      61.32$3.3215UnscoredUnpriced
                      Grok 4.6 · high
                      Missing evidence
                        • Execute · DeepSWE: The Coding Agent Index does not publish this model
                        • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                        • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                        60.80$2.5600UnscoredUnpriced
                        Muse Spark 1.3 · xhigh59.25$1.752264.18$1.6150
                        Grok 4.6 · medium
                        Missing evidence
                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                          58.85$1.8538UnscoredUnpriced
                          GPT-5.6 Sol · max58.50$2.136065.05$4.9955
                          GLM-5.3-Flash · max
                          Missing evidence
                            • Execute · DeepSWE: The Coding Agent Index does not publish this model
                            • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                            • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                            57.70$0.4414UnscoredUnpriced
                            GPT-5.6 Sol · xhigh56.34$1.634063.34$3.7367
                            GPT-6 Astra · low
                            Missing evidence
                              • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                              • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                              • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                              55.91$0.8875UnscoredUnpriced
                              Qwen3.8 27B · xhigh
                              Missing evidence
                                • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                54.08$1.1684UnscoredUnpriced
                                Muse Spark 1.2 · xhigh53.76$0.989561.64$2.0728
                                Claude Sonnet 5 · max
                                Missing evidence
                                  • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                  • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                  • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                  53.21$7.4509UnscoredUnpriced
                                  GPT-5.6 Sol · high53.20$1.112864.12$3.0005
                                  Claude Opus 5 · low52.24$1.023459.40$2.2952
                                  Claude Opus 4.8 · max51.91$4.405362.14$7.7247
                                  Grok 4.5 · high51.66$0.670064.09$2.4390
                                  DeepSeek V4 Pro 0813 · max51.21$0.545642.78$0.0908
                                  Grok 4.6 · low
                                  Missing evidence
                                    • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                    • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                    • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                    50.87$0.4568UnscoredUnpriced
                                    GPT-5.6 Terra · max50.57$1.499160.42$1.9300
                                    GPT-5.6 Terra · xhigh49.56$0.958056.03$1.3589
                                    GPT-5.6 Luna · max49.52$0.207757.17$0.2883
                                    GPT-5.5 · xhigh49.01$2.748861.03$4.7524
                                    GPT-5.6 Sol · medium48.61$0.556861.62$2.1930
                                    Gemini 3.8 Flash · high48.60$1.384961.15$2.0378
                                    DeepSeek V4 Flash 0731 · max48.46$0.230149.76$0.0595
                                    GPT-5.5 · high
                                    Missing evidence
                                      • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                      • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                      • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                      47.30$1.7874UnscoredUnpriced
                                      GPT-5.6 Luna · xhigh46.31$0.113752.96$0.2359
                                      Gemini 3.7 Flash · high45.28$1.099159.63$1.2679
                                      GPT-5.6 Terra · high44.37$0.474654.65$1.1409
                                      Muse Spark 1.1 · xhigh43.88$1.301154.92$1.4352
                                      GPT-5.5 · medium43.23$1.071055.31$2.6525
                                      GPT-5.6 Luna · high42.45$0.063651.69$0.1798
                                      GPT-5.6 Sol · low41.36$0.277655.23$1.2925
                                      Gemini 3.8 Flash · low
                                      Missing evidence
                                      • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                      • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                      • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                      • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                      40.70UnpricedUnscoredUnpriced
                                      GPT-5.6 Terra · medium37.70$0.229148.05$0.6700
                                      GPT-5.6 Terra · low35.00$0.184538.65$0.3851
                                      Gemini 3.5 Flash · high
                                      Missing evidence
                                        • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                        • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                        • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                        34.72$2.0846UnscoredUnpriced
                                        Gemini 3.5 Flash · medium
                                        Missing evidence
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          34.70$2.4744UnscoredUnpriced
                                          GPT-5.6 Luna · medium31.84$0.018541.97$0.0884
                                          GPT-5.6 Luna · low24.32$0.009825.06$0.0389
                                          Gemini 3.7 Flash · low
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                          UnscoredUnpricedUnscoredUnpriced
                                          Gemini 3.7 Flash · medium
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                          UnscoredUnpricedUnscoredUnpriced
                                          Gemini 3.8 Flash · medium
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                          UnscoredUnpricedUnscoredUnpriced
                                          Kimi K3 · low
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                          UnscoredUnpricedUnscoredUnpriced
                                          Qwen3.8-Flash-Next · xhigh
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          Agnes 2.5 Pro Beta · reasoning
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          GPT-5.5 · low
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index has not run this model at this reasoning setting
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index has not run this model at this reasoning setting
                                          UnscoredUnpricedUnscoredUnpriced
                                          Qwen3.8 27B · low
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          Qwen3.8 27B · medium
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          GPT-5.4 Pro · xhigh
                                          Missing evidence
                                          • Figure · HLE: No published measurement at this exact configuration
                                          • Figure · HLE: Direct task accounting is unavailable
                                          • Figure · AA-LCR 1.1: No published measurement at this exact configuration
                                          • Figure · AA-LCR 1.1: Direct task accounting is unavailable
                                          • Figure · Omniscience reliability: No published measurement at this exact configuration
                                          • Figure · Omniscience reliability: Direct task accounting is unavailable
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          GPT-5.5 Pro · xhigh
                                          Missing evidence
                                          • Figure · HLE: No published measurement at this exact configuration
                                          • Figure · HLE: Direct task accounting is unavailable
                                          • Figure · AA-LCR 1.1: No published measurement at this exact configuration
                                          • Figure · AA-LCR 1.1: Direct task accounting is unavailable
                                          • Figure · Omniscience reliability: No published measurement at this exact configuration
                                          • Figure · Omniscience reliability: Direct task accounting is unavailable
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          Gemini 3.5 Flash · minimal
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          DeepSeek V4 Flash Vision · max
                                          Missing evidence
                                          • Figure · Briefcase analytical quality: No published measurement at this exact configuration
                                          • Figure · Briefcase analytical quality: Direct task accounting is unavailable
                                          • Execute · DeepSWE: The Coding Agent Index does not publish this model
                                          • Execute · Terminal-Bench 2.1: The Coding Agent Index does not publish this model
                                          • Execute · SWE-Atlas-QnA: The Coding Agent Index does not publish this model
                                          UnscoredUnpricedUnscoredUnpriced
                                          Claude Fable 5 · high
                                          Missing evidence
                                          • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                          • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                          • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                          • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                            UnscoredUnpriced65.13$5.9710
                                            Claude Fable 5 · low
                                            Missing evidence
                                            • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                            • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                            • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                            • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                              UnscoredUnpriced58.56$3.1661
                                              Claude Fable 5 · medium
                                              Missing evidence
                                              • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                              • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                              • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                              • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                UnscoredUnpriced63.20$4.7355
                                                Claude Fable 5 · xhigh
                                                Missing evidence
                                                • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                  UnscoredUnpriced66.02$8.5258
                                                  Kimi K3 · unspecified
                                                  Missing evidence
                                                  • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                  • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                  • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                  • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                    UnscoredUnpriced62.64$3.0816
                                                    GPT-5.6 Luna · none
                                                    Missing evidence
                                                    • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                    • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                    • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                    • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                      UnscoredUnpriced19.10$0.0695
                                                      Claude Opus 4.8 · high
                                                      Missing evidence
                                                      • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                      • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                      • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                      • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                        UnscoredUnpriced57.58$3.7823
                                                        Claude Opus 4.8 · low
                                                        Missing evidence
                                                        • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                        • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                        • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                        • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                          UnscoredUnpriced49.04$2.1826
                                                          Claude Opus 4.8 · medium
                                                          Missing evidence
                                                          • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                          • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                          • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                          • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                            UnscoredUnpriced55.77$3.3038
                                                            Claude Opus 4.8 · xhigh
                                                            Missing evidence
                                                            • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                            • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                            • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                            • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                              UnscoredUnpriced59.07$5.6654
                                                              Claude Opus 5 · none
                                                              Missing evidence
                                                              • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                              • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                              • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                              • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                                UnscoredUnpriced59.17$3.5309
                                                                Qwen3.8 Max · unspecified
                                                                Missing evidence
                                                                • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                                • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                                • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                                • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                                  UnscoredUnpriced61.31$3.2306
                                                                  GPT-5.6 Sol · none
                                                                  Missing evidence
                                                                  • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                                  • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                                  • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                                  • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                                    UnscoredUnpriced43.37$1.0869
                                                                    GPT-5.6 Terra · none
                                                                    Missing evidence
                                                                    • Figure · HLE: Artificial Analysis has not published a model row for this configuration
                                                                    • Figure · AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
                                                                    • Figure · Omniscience reliability: Artificial Analysis has not published a model row for this configuration
                                                                    • Figure · Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
                                                                      UnscoredUnpriced23.09$0.2927
                                                                      Data: September 6, 2026 · Artificial Analysis, Cognition, ZapierBenchmark API costs compare setups. They are not your bill.