Which model, at which effort?
Depends what the whole job is worth to you. Drag the line.
Figure it out. The highest measured score for mixed investigation. Every cheaper rung gives something up.
Execute it. The highest measured score for coding agent work. Every cheaper rung gives something up.
Every measured configurationquality against cost · open ↓
See the trade-off
Mixed investigation · higher is better, left is cheaper
Measured frontiers
Claude Fable 5.1 · maxSelected · 72.68 points · $11.93 / task
Chart key & sources
Models you can use33 of 33 included · open ↓
Turn off a provider or model you cannot use. The rail recomputes from what remains.
Select a provider or model to include it. Open effort settings to exclude individual configurations.
5/5 efforts · edit
7/7 efforts · edit
7/7 efforts · edit
7/7 efforts · edit
4/4 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
5/5 efforts · edit
6/6 efforts · edit
5/5 efforts · edit
5/5 efforts · edit
5/5 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
2/2 efforts · edit
2/2 efforts · edit
4/4 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
3/3 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
2/2 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
1/1 efforts · edit
How this is measuredsources, weights and limits · open ↓
How the rail is builtundominated rungs on a log job-share axis
The rail measures a whole job, not one benchmark task. Each dot is a model and effort that no cheaper measured configuration beats on the profile score. A dot sits at what that rung contributes to the job: its lane's share of the job, times what the rung costs against the best rung in its own lane. The right end of the rail is the job with both lanes at their best, so the ticks read 1, .1, .01 — fractions of that job, not dollars. Equal distances still mean equal cost ratios.
Why not one dollar figure for both lanes? Because a benchmark task is not a job, and the two lanes' tasks are not the same size. A Figure task is a set of hard questions and a business analysis, scored per question. An Execute task is a repository issue or a terminal session an agent works for minutes. Charging both lanes the same dollar ceiling treats one of each as an equal purchase — which it is not, when a real job runs one lane once and the other for hours.
"Your job is" sets the split. Mostly figuring puts 20% of the job in Execute, Balanced puts half, Mostly executing puts 80% there. The line is then one threshold both lanes' contributions sit under, and it converts back to a dollar ceiling per lane before anything is chosen — the lane carrying most of the job gets squeezed first, which is the point. Each answer still shows its own benchmark dollars per task and the date of the capture it came from.
Those three numbers are chosen, not measured. They are round positions, from the author's own working shape averaged informally with colleagues'. There is no retained source behind 20, 50 and 80, and none is claimed: they exist so a reader can say roughly what their week looks like, and a spinner would promise a precision nobody here has. Pick the one nearest your own and read the answer it gives.
The author's own jobs are a separate thing, and they are measured. Across 16 of his Claude Code jobs that ran both a figuring phase and an executing phase, 2026-09-03 to 09-08, executing took 5 to 94 percent of those two phases' API-equivalent cost at that day's published rates — median 84 percent, 75 percent weighted by dollars, over $347 of attributed spend. The largest such job spent $63 of $77 executing. That is one person's five days on one harness. It is not a bill, it is not a measurement of your work, and the width of that range is exactly why the split is a control rather than a constant.
Those runs were also almost all cache reads: 98.4 percent of input-side tokens across the Claude Code sessions, and 96.1 percent of input tokens across 258 Codex sessions in the same window. For work shaped like that a model's cache-read price orders the bill far more than its input price does, and a benchmark task bounded at a few minutes cannot show you that. Opus 5 reads cache at $0.50 per million tokens, Fable 5.1 at $0.25, Sonnet 5 at $0.20, Haiku 4.5 at $0.10 (Anthropic pricing, read 2026-09-08).
The rail carries the undominated rungs only. Cheaper and stronger steps are the dots themselves: every measured Pareto step for both lanes is on the line, and dragging to one selects it. A ceiling below every rung produces no recommendation for that lane and names the cheapest measured option instead.
How alternatives qualifyscreening rules and evidence limits
A model does not need every benchmark to earn consideration. One strong, relevant evaluation can qualify a scoped alternative. Figure looks across practical analysis, reasoning, factual reliability and document work. Execute looks across coding-agent work, professional outputs and tool workflows. A specialist remains a specialist; its best result does not establish broad capability.
Within your available models and effort settings, each benchmark admits results within 5 score points of its leader, or 100 Elo on native Elo boards. It also admits undominated quality–cost options within 20 points, or 400 Elo, of the leader. “Undominated” means no measured option is both at least as good and at least as cheap, with one strict advantage. Missing cost cannot earn a value claim. These margins are chosen screening rules, not uncertainty estimates or proof of a meaningful lead. Cheaper results below the quality floor remain in the dated discovery survey for follow-up.
The selected focus shows its core benchmarks and relevant supporting evidence. Coding agent work also shows FrontierCode, Zapier, AA's own Terminal-Bench run and SciCode; Deliverables shows Briefcase rubric checks and agentic professional work; Codebase investigation shows long-context reasoning and document reading. Supporting scores do not enter the profile average. Every qualifying model remains accessible; the first six are a compact alphabetical preview. Open a model’s row to see the exact efforts, benchmark scores, source, date, harness and fallback status. Several strong results from one publisher are not independent corroboration. Different effort settings on one card are separate observations, not a combined configuration.
The shortlist uses native Briefcase and GDPval Elo and the published AA-Omniscience index. Composite profiles use the fixed conversions below; the native evidence stays visible beside them. Native and composite units must not be compared directly. FrontierCode Fable runs have undisclosed fallback settings; Kimi's Zapier effort is undisclosed. These results can support scoped consideration, but cannot fill an exact broad-composite gap.
How the recommendation is pickedquality, spending and uncertainty
Start with quality, then decide how much to spend. Quality first selects the highest measured profile score among configurations with all required quality evidence, comparable API costs and reproducible settings. Equal scores prefer lower cost, then a stable configuration identity. It does not exchange score points for a guessed dollar value, and it does not claim to identify the best unmeasured model.
Choose Figure it out or Execute it, then narrow the work when a focused profile fits. A deliverables result can qualify on its own evidence even when the Coding Agent Index has not run that configuration and so cannot give it a broad Execute score. Undisclosed effort or fallback settings remain scoped observations; they cannot set the recommendation.
The rail sets a logarithmic ceiling in shares of one job: equal distances mean equal cost ratios, and the position converts to a dollar ceiling per lane through that lane's share of the job. Each kind of work then answers with its highest score within its own ceiling. Every measured Pareto step is a dot on the rail, so dragging past one selects it. A ceiling too low produces an explicit empty recommendation with the cheapest available alternative. Releasing the rail at its top removes the ceiling — Quality first — without changing model or effort exclusions.
Cheaper models that are not on the frontier still appear where their evidence earns it: the shortlist under Why this choice states each qualifying model’s measured scope, system, and unknown or above-ceiling costs. A component strength is not a broad recommendation.
The audit tests the policy, not the probability of being right. Open Method sensitivity & winning margins to compare half/double Elo slopes, wider group weights, omitted benchmarks/groups and focused compositions. Fixed cohorts keep the baseline configurations; expanded cohorts separately admit newly scoreable ones. Quality-only results allow unknown costs; priced results require comparable costs and apply your ceiling. Your access filters apply throughout. Scenario counts are not confidence percentages.
The reproducible all-model audit retains winners, coverage and native component margins. A stable result can still be wrong for your task: the scenarios cannot supply representative real-world outcomes that were never measured.
Why this choicemeasurements, method & access
55 of 83 evaluated configurations can set this profile’s recommendation.
Why 28 configurations cannot set this recommendation
- Gemini 3.7 Flash low · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Gemini 3.7 Flash medium · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Gemini 3.8 Flash low · AA benchmark harness
40.70 points · Comparable cost unavailable
- Briefcase analytical quality: Direct task accounting is unavailable
- Gemini 3.8 Flash medium · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Kimi K3 low · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Qwen3.8-Flash-Next xhigh · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Agnes 2.5 Pro Beta reasoning · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- GPT-5.5 low · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Qwen3.8 27B low · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Qwen3.8 27B medium · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- GPT-5.4 Pro xhigh · AA benchmark harness
Score unavailable · Comparable cost unavailable
- HLE: No published measurement at this exact configuration
- HLE: Direct task accounting is unavailable
- AA-LCR 1.1: No published measurement at this exact configuration
- AA-LCR 1.1: Direct task accounting is unavailable
- AA-Omniscience index: No published measurement at this exact configuration
- AA-Omniscience index: Direct task accounting is unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- GPT-5.5 Pro xhigh · AA benchmark harness
Score unavailable · Comparable cost unavailable
- HLE: No published measurement at this exact configuration
- HLE: Direct task accounting is unavailable
- AA-LCR 1.1: No published measurement at this exact configuration
- AA-LCR 1.1: Direct task accounting is unavailable
- AA-Omniscience index: No published measurement at this exact configuration
- AA-Omniscience index: Direct task accounting is unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Gemini 3.5 Flash minimal · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- DeepSeek V4 Flash Vision max · AA benchmark harness
Score unavailable · Comparable cost unavailable
- Briefcase analytical quality: No published measurement at this exact configuration
- Briefcase analytical quality: Direct task accounting is unavailable
- Claude Fable 5 high · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Fable 5 low · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Fable 5 medium · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Fable 5 xhigh · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Kimi K3 unspecified · Not published for this configuration
Score unavailable · Comparable cost unavailable
Reasoning effort is not disclosed; cannot reproduce an exact setting
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- GPT-5.6 Luna none · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Opus 4.8 high · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Opus 4.8 low · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Opus 4.8 medium · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Opus 4.8 xhigh · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Claude Opus 5 none · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Qwen3.8 Max unspecified · Not published for this configuration
Score unavailable · Comparable cost unavailable
Reasoning effort is not disclosed; cannot reproduce an exact setting
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- GPT-5.6 Sol none · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- GPT-5.6 Terra none · Not published for this configuration
Score unavailable · Comparable cost unavailable
- HLE: Artificial Analysis has not published a model row for this configuration
- HLE: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-LCR 1.1: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- AA-Omniscience index: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
- Briefcase analytical quality: Artificial Analysis has not published a model row for this configuration
What the chosen configuration was measured on
The weights are declared policy, not a validated prediction of your work. Collaborative framing and planning remain unmeasured.
Score & native measurements
72.68 points · Mixed investigation. Benchmark proxies do not guarantee correctness.
| Benchmark | Native result | API $ / task |
|---|---|---|
| HLE16.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 59.13% | $1.59 |
| AA-LCR 1.116.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 85.33% | $1.58 |
| AA-Omniscience index16.7% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 43.45 points | $0.27 |
| Briefcase analytical quality50.0% weight · AA benchmark harnessClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 1965.99 Elo | $22.71 |
Supporting evidence, outside the score: Briefcase presentation 1475.42 Elo · 2026-09-06. Same exact model and effort; AA benchmark harness.
Access: subscription, API or your hardware
Reviewed 2026-09-06. Availability does not reproduce a benchmark’s agent, effort or fallback configuration.
Subscription Verified route
Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start.
Plan-dependent inclusion or credits; the 50% is not extra capacity. Usage-based Enterprise uses API rates.
Claude Code requires 2.1.255 or newer. App fallback is automatic; availability alone does not reproduce the benchmark fallback system.
API Verified route
claude-fable-5-1
$10 input / $50 output USD per million tokens.
API fallback must be configured. Cost and quality depend on the exact fallback model/settings, not only the primary model. Effort: low through max; check native menu.
Self-hosting Not verified
No verified open-weight or self-hosting route for this exact model.
Not verified
Do not infer exact-version availability or included usage from a family name.
Method sensitivity & winning margins
These are declared scenarios, not confidence percentages. They vary Elo slope by half/double, group weights by 0.5/1/2, omit benchmarks/groups and test focused compositions. Your setup and exclusions apply. Priced results apply your spending ceiling; quality-only results ignore costs.
Fixed cohorts keep the baseline eligible configurations; expanded cohorts admit configurations newly scoreable under each composition. Neither fills missing scores. Several components share Artificial Analysis, model judges and agent tools; they are not independent confirmations. Benchmark-specific harness and fallback limits remain.
Baseline: 55 priced configurations from 24 models. Claude Fable 5.1 max leads Claude Fable 5.1 xhigh by 1.308 profile units.
- HLE: native margin +0.0042 at 16.7% weight.
- AA-LCR 1.1: native margin +0.0233 at 16.7% weight.
- AA-Omniscience index: native margin +1.0667 at 16.7% weight.
- Briefcase analytical quality: native margin +30.4100 at 50.0% weight.
| Scenario | Fixed · quality only | Fixed · priced | Expanded · quality only | Expanded · priced |
|---|---|---|---|---|
| Declared profile | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Elo slope 0.5× | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Elo slope 2× | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×0.5, foundation ×0.5 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×0.5, foundation ×1 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×0.5, foundation ×2 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×1, foundation ×0.5 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×1, foundation ×1 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×1, foundation ×2 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×2, foundation ×0.5 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×2, foundation ×1 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Group weights analysis ×2, foundation ×2 | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Without hle | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Without lcr | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Without omniscience | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Without briefcaseAnalysis | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max67 configurations · 27 models | Claude Fable 5.1 max67 configurations · 27 models |
| Without analysis | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max67 configurations · 27 models | Claude Fable 5.1 max67 configurations · 27 models |
| Without foundation | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models | Claude Fable 5.1 max56 configurations · 24 models | Claude Fable 5.1 max55 configurations · 24 models |
| Codebase investigation alone | No comparable result0 configurations · 0 models | No comparable result0 configurations · 0 models | Claude Fable 5.1 max47 configurations · 17 models | Claude Fable 5.1 max47 configurations · 17 models |
Evidence for mixed investigation
19 modelsRelevant benchmark strengths, including setups that cannot set a priced recommendation. Supporting benchmarks stay separate from the profile score.
Agnes 2.5 Pro BetaLong-context reasoningMeasured with limits · One publisher
These strengths do not establish broad capability.
- AA-LCR 1.1 ↗reasoning effort · 83.00 points · $0.012 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Claude Fable 5Scientific reasoning · Hard reasoning +1Comparable profile option · One publisher
At least one qualifying effort has complete, priced profile evidence.
Scientific reasoning · Hard reasoning · Factual reliability
- CritPt ↗max effort · 28.57 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- HLE ↗max effort · 55.47 points · $0.93 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-Omniscience index ↗max effort · 43.30 points · $0.070 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
Claude Fable 5.1Practical analysis · Scientific reasoning +3Comparable profile option · One publisher
At least one qualifying effort has complete, priced profile evidence.
Practical analysis · Scientific reasoning · Hard reasoning · Long-context reasoning · Factual reliability
- Briefcase analytical quality ↗max effort · 1965.99 Elo · $22.71 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- Briefcase analytical quality ↗xhigh effort · 1935.58 Elo · $17.79 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- Briefcase analytical quality ↗high effort · 1847.00 Elo · $11.40 / benchmark taskQuality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- CritPt ↗max effort · 29.71 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- CritPt ↗xhigh effort · 31.14 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- CritPt ↗high effort · 30.29 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- CritPt ↗medium effort · 29.14 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- CritPt ↗low effort · 27.71 points · cost unavailableNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- HLE ↗max effort · 59.13 points · $1.59 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- HLE ↗xhigh effort · 58.71 points · $1.03 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- HLE ↗high effort · 55.93 points · $0.40 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- HLE ↗medium effort · 53.80 points · $0.23 / benchmark taskQuality–cost frontier · AA benchmark harness · fallback enabled · 2026-09-06
- AA-LCR 1.1 ↗max effort · 85.33 points · $1.58 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-LCR 1.1 ↗high effort · 83.67 points · $1.48 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-LCR 1.1 ↗medium effort · 84.67 points · $1.48 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-Omniscience index ↗max effort · 43.45 points · $0.27 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-Omniscience index ↗xhigh effort · 42.38 points · $0.077 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
- AA-Omniscience index ↗high effort · 40.80 points · $0.021 / benchmark taskNear the available leader · AA benchmark harness · fallback enabled · 2026-09-06
Claude Opus 5Practical analysis · Scientific reasoning +2Comparable profile option · One publisher
At least one qualifying effort has complete, priced profile evidence.
Practical analysis · Scientific reasoning · Hard reasoning · Factual reliability
- Briefcase analytical quality ↗max effort · 1909.74 Elo · $17.79 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
- Briefcase analytical quality ↗xhigh effort · 1900.45 Elo · $14.26 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
- Briefcase analytical quality ↗high effort · 1815.72 Elo · $10.41 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
- CritPt ↗max effort · 29.14 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
- CritPt ↗xhigh effort · 27.71 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
- CritPt ↗high effort · 28.29 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
- HLE ↗max effort · 54.87 points · $0.67 / benchmark taskNear the available leader · AA benchmark harness · single model · 2026-09-06
- HLE ↗xhigh effort · 54.40 points · $0.52 / benchmark taskNear the available leader · AA benchmark harness · single model · 2026-09-06
- AA-Omniscience index ↗medium effort · 31.02 points · $0.0062 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Gemini 3.7 FlashFinancial analysis · Hard reasoningComparable profile option · One publisher
At least one qualifying effort has complete, priced profile evidence.
- AA-AnalystAgent pass@1 ↗high effort · 70.50 points · cost unavailableNear the available leader · AA benchmark harness · single model · 2026-09-06
- HLE ↗high effort · 47.87 points · $0.037 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
Gemini 3.8 FlashHard reasoning · Long-context reasoning +1Measured with limits · One publisher
These strengths do not establish broad capability.
Hard reasoning · Long-context reasoning · Factual reliability
- HLE ↗medium effort · 42.12 points · $0.025 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
- AA-LCR 1.1 ↗medium effort · 84.00 points · $0.090 / benchmark taskNear the available leader · Quality–cost frontier · AA benchmark harness · single model · 2026-09-06
- AA-Omniscience index ↗medium effort · 28.58 points · $0.0036 / benchmark taskQuality–cost frontier · AA benchmark harness · single model · 2026-09-06
How options qualify & what this search misses
Within your available configurations, each benchmark admits results within 5 points of its leader (100 for Elo), plus quality–cost frontier results within 20 points (400 Elo). These are screening choices, not confidence intervals. No cost means no value claim. Known weak results remain available in each source; a strength is not an endorsement of the whole model.
The 2026-09-06 survey examined 643 AA configurations, 95 FrontierCode entries and 95 Zapier entries. Coverage and public availability remain incomplete; this is a source-backed shortlist, not proof that every great model was found.
Access reference for all 33 maintained models
Access: subscription, API or your hardware
Reviewed 2026-09-06. Availability does not reproduce a benchmark’s agent, effort or fallback configuration.
Subscription Not verified
Vendor advertises a Token Plan; exact inclusion and pricing were not verified.
Not verified
Do not infer exact-version availability or included usage from a family name.
API Verified route
First-party Pro Beta API is announced by the vendor.
$0.10 input / $0.30 output per million tokens; cached input $0.01.
Pricing evidence is a vendor social post, not a verified API tariff page. Beta status remains; configurable reasoning efforts unverified.
Self-hosting Not verified
No verified open-weight or self-hosting route for this exact model.
Not verified
Do not infer exact-version availability or included usage from a family name.
Use your allowance wellchat, agent and when to switch
Practical guidance for each kind of work, from the configuration the rail names. There is no measured quota optimum.
Use your allowance well.
Settle the decision
Decide what success means with the model that has the useful context.
Your agentClaude Fable 5.1 · max
Start here when the investigation needs files, tests or tools.
Limits & when to switch
On Claude Pro and Max, Chat and Claude Code share usage limits. Switching between them adds no separate allowance.
Skip a separate chat when the plan is settled or the decision needs repository tools. Reconsider the model or effort when the design is unresolved or fixes keep failing; carry the useful context with you.
Use Cheaper to inspect the measured trade-off. API savings do not measure subscription savings, and the benchmarks do not identify when to escalate.
Selected model access: Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start. Check access & evidence below.
Product facts checked 2026-09-07. Claude Pro and Max usage
Practical guidance · no measured quota optimum.
Use your allowance well.
Settle the decision
Decide what success means with the model that has the useful context.
Claude CodeClaude Fable 5.1 · max
Bring the decision. Build, verify and fix in the same task.
Limits & when to switch
On Claude Pro and Max, Chat and Claude Code share usage limits. Switching between them adds no separate allowance.
Skip a separate chat when the plan is settled or the decision needs repository tools. Reconsider the model or effort when the design is unresolved or fixes keep failing; carry the useful context with you.
Use Cheaper to inspect the measured trade-off. API savings do not measure subscription savings, and the benchmarks do not identify when to escalate.
Selected model access: Max, premium Team and premium legacy seat-based Enterprise seats include Fable within up to 50% of the shared weekly allowance. Pro, standard Team and standard Enterprise seats use purchased usage credits from the start. Check access & evidence below.
Product facts checked 2026-09-07. Claude Pro and Max usage
Practical guidance · no measured quota optimum.
Subscription, API or self-hostingwhat the cost numbers mean
Benchmark API cost is a comparison unit, not your bill. Subscription plans may include an allowance, require credits from the first request, restrict supported tools or leave exact versions undisclosed. The selected model’s access disclosure and the full access reference keep those distinctions separate, with official sources and verification dates.
Downloadable weights establish a self-hosting route, not laptop suitability or savings. The reference gives sourced serving examples and distinguishes author formats, community quantizations and related hosted variants. For example, hosted Qwen Flash is based on Flash-Next with additional production features; it is not silently treated as the same measured weights configuration. Local hardware, electricity and operations costs are not estimated here.
What turning the effort dial up costsmeasured steps and regressions
Turning the dial up buys the model more thinking. You pay for that thinking in tokens, so it is a spending dial first and a quality dial second.
That makes it measurable. Every row here is one model at one setting, scored on DeepSWE, Terminal-Bench 2.1 and SWE-Atlas-QnA and billed with the same job weights. The ladder compares complete measured settings; a missing setting is skipped, not estimated.
GPT-5.6 Sol shows it most starkly. Among models with at least four complete settings, the last step costs about 44× what the first did: $0.017 a point at the bottom, $0.74 at the top. Measured September 6, 2026.
| Step | Score | Points bought | Added cost | $ / point |
|---|---|---|---|---|
| none → low | 55.23 | +11.86 | $0.2056 | $0.0173 |
| low → medium | 61.62 | +6.39 | $0.9005 | $0.1409 |
| medium → high | 64.12 | +2.50 | $0.8075 | $0.3230 |
| high → xhigh | 63.34 | −0.78 | $0.7362 | bought nothing |
| xhigh → max | 65.05 | +1.71 | $1.2588 | $0.7361 |
No number here tells you where to stop. The board prices a point at each step. What a wrong answer costs you is the other half, and only you have it.
Higher does not always mean better, either. 2 measured configurations score lower at a higher setting:
- Claude Opus 5 · xhigh → max — 1.12 points down, $0.77 more per task. Paying more for a worse result.
- GPT-5.6 Sol · high → xhigh — 0.78 points down, $0.74 more per task. Paying more for a worse result.
One board, one capture date, no error bars. Read that as a reason to check your own setting, not as a law about the model. Either way, “turn it to max” is a guess rather than advice.
Never pick the setting on its own. A cheap model turned up and a costly one turned down are the same kind of thing here — which is why the card compares configurations rather than ranking model names.
What each task is measured onsources, weights and missing evidence
A profile defines the work being compared. Four options, two per lane, each scored by one named source — so a focus can never quietly mix in evidence its own definition does not name. There used to be eight, which asked you which benchmark you trusted rather than what you were about to do.
| Lane | Profile | Required evidence and weights |
|---|---|---|
| Figure it out | Mixed investigation | Half reasoning/knowledge foundation, half Briefcase analytical quality — the AA model row, in points |
| Figure it out | Codebase investigation | SWE-Atlas-QnA alone, from the Coding Agent Index row for your harness, in percent correct |
| Execute it | Coding agent work | The AA Coding Agent Index for your harness: DeepSWE, Terminal-Bench 2.1 and SWE-Atlas-QnA in equal thirds, in points |
| Execute it | Deliverables | GDPval-AA v2 alone, in native Elo |
What scores the Execute lane
The Execute lane reads the Artificial Analysis Coding Agent Index, captured at v1.4 on 2026-09-08, whose own page states it "incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA" in equal thirds, over three attempts per task.
A row on that board is a harness, a model and a reasoning setting together — "Claude Code / Fable 5.1 / max", not "Fable 5.1". That is the whole reason the lane moved to it: the number describes the pair you actually run, not a model in a scaffold nobody uses. Choosing Claude Code reads Claude Code's rows; choosing Codex reads Codex's; All models takes each configuration's best published row whatever agent produced it, and the answer names that agent. Cost is that row's own mean API dollars per task, cached input and cache writes included.
Its three parts: DeepSWE resolves repository issues end to end (113 tasks); Terminal-Bench v2.1 runs command-line coding, administration and data tasks (89 tasks); SWE-Atlas-QnA answers expert-written comprehension questions inside a repository (124 tasks). The last of those, alone and in percent, is what Codebase investigation scores on the Figure lane.
HLE tests hard academic questions; AA-LCR tests reasoning across document sets. Omniscience supplies the publisher's factual-reliability score. Briefcase separates analytical quality, rubric checks and presentation; its presentation score is never substituted for analysis. GDPval-AA v2 compares professional deliverables and is what Deliverables scores, alone, in its native Elo. FrontierCode Main and Zapier AutomationBench are now supporting evidence on the rows that carry them, not requirements: neither gates a lane any more.
Fixed conversions preserve differences without depending on the current roster. Composite profiles use:
Rate p → 100 × p
Published Omniscience O → (O + 100) / 2
Native Briefcase/GDPval E → (E − 500) / 20Elo is not clipped: 1000 becomes 25 points, 1500 becomes 50 and 2500 becomes 100; higher or negative values remain possible. Twenty Elo per point is a declared cross-benchmark policy, not an empirically established exchange rate. 1000 Elo is not asserted to mean human-level performance. Single-benchmark profiles keep their native units. Adding or filtering models cannot move another configuration's score.
Costs use the same profile weights as quality. AA bills use per-benchmark input/output tokens, cache rates, prices and fallback fractions, divided by task count. Both Briefcase dimensions use its full task bill. The Coding Agent Index publishes its own mean API dollars per task per row, already inclusive of cached input and cache writes. FrontierCode publishes mean API dollars per rollout, already per task. Zapier publishes API dollars per task rounded to cents. Dedicated-deployment prices and incomplete fallback bills remain unpriced. Subscription inclusion is never substituted with zero benchmark cost.
Each profile requires all its components at the exact model, effort, version, harness and compatible system. Missing values stay unscored or unpriced; weights are never renormalized around them. A Coding Agent Index row is used only for the harness that produced it — a Codex number never wears a Claude Code label. A fallback-enabled AA system cannot borrow a standalone external result, even when observed fallback use is zero; the index's own fallback-enabled rows stay eligible and say so, the same rule the Figure lane applies to AA's. Focused source joins preserve model, effort, harness, fallback identity and source record. An estimated overall AA Index neither fills a gap nor excludes complete evidence.
What this composition costs you
Coverage. The index publishes 68 rows, of which 53 name a model compared here; they cover 49 model-and-effort configurations as of 2026-09-08 — 17 reachable under Claude Code and 21 under Codex once model compatibility is applied. A configuration the board has not run carries no Execute score, with that reason attached, and stays visible in the shortlist. Claude Fable 5.1 is published only at max, in Claude Code; GPT-6 Astra only at max, in Codex — so a reader running either below max is answered from the shortlist instead. The composition this one replaced required FrontierCode, Briefcase rubric, GDPval and Zapier all present, which admitted 29 configurations and excluded every Fable row — the reader's own model missing from the lane that decides their biggest bill.
The demoted sources still agree with it. Over the 31 configurations both measure, FrontierCode ranks them at Spearman 0.90 against the index; Zapier, over 28, at 0.89. Requiring FrontierCode was buying agreement the index already had, at the price of every configuration FrontierCode does not publish. Both remain on the rows.
Where the ordering comes from. Inside the index, Terminal-Bench 2.1 sits near its ceiling for the leading rows: across the ten highest-scoring rows in the 2026-09-08 capture it averages 86.6 against a board maximum of 91.0, while DeepSWE averages 65.0 and SWE-Atlas-QnA 49.2. The top of the board is therefore separated mostly by DeepSWE and SWE-Atlas-QnA, not by the terminal component.
The board moves. Between May and August 2026 the index went through four versions. Its benchmark line-up changed once, at v1.1 in June 2026, which added DeepSWE and removed SWE-Bench-Pro-Hard-AA. v1.2 and v1.3 changed how SWE-Atlas-QnA is scored — rubric reward, then binary pass/fail, then Scale's own Task Resolve Rate — and v1.4 upgraded Terminal-Bench from v2 to v2.1 and added reward-hacking detection (AA methodology, read 2026-09-08). Scores are not comparable across those versions. The extractor asserts the published version sentence, the exact set of dataset slugs and the equal-third weights, and stops the refresh rather than re-weighting the lane underneath you if any of them changes.
Its two coding parts each carry a caveat. DeepSWE (Datacurve, July 2026) is 113 original tasks over 91 repositories, graded binary with no partial credit, run for every model in one fixed harness — a deliberate trade its own paper describes as removing the scaffolding confound at the cost of cross-harness realism. It is also a single-vendor benchmark whose contamination controls cannot be audited from outside. SWE-Atlas-QnA (Scale AI, March 2026) is 124 questions over 11 production repositories; Scale's own leaderboard notes some configurations show an elevated refusal rate because the tasks trip security filters, which reads as a low score without being one.
Harness effects are the point, and also a limit. The number describes one agent at one version, and it does not transfer to the same model in another scaffold. That is not a small correction: a controlled 3×3 study on a 100-task SWE-bench Verified subset measured harness variance at 7.8× model variance and found rank reversals in 6 of 9 model-pair comparisons (arXiv 2605.23950, a position paper whose authors explicitly decline to claim that ratio is universal), and Harness-Bench found a 23.8-point aggregate gap between the best and worst configurable harnesses over the same 106 tasks and model pool (arXiv 2605.27922). Both were read 2026-09-08. This is the argument for scoring the pair you run — and the reason a number from one agent tells you little about another.
What the Figure lane's own sources are worth
The Execute lane is not the only one holding evidence with a record. Stated plainly, and dated:
- HLE has contested answers. FutureHouse examined 321 text-only chemistry and biology questions and found that 29% ± 3.7 had answers contradicted by peer-reviewed literature (FutureHouse, 23 July 2025). The HLE team's own follow-up on that page put roughly 18% of a reviewed subset as problematic under three-expert review and committed to rolling revisions. The finding covers that subset, not the whole exam.
- AA-Omniscience's questions are machine-written. Artificial Analysis states its 6,000 questions were generated by an LLM-based question-generation agent from authoritative academic and industry sources (AA, read 2026-09-08). The sources are human; the questions are not.
- AA-LCR v1.1 corrected its own answers. AA's methodology records that v1.1 added a grading system prompt, corrected 16 answer keys and changed the grader, and that scores are not directly comparable with v1.0 (AA methodology, read 2026-09-08).
- Briefcase is Artificial Analysis's own private benchmark. Its 91 tasks across four held-out scenarios were written by outside experts from firms including Google, McKinsey and BCG, but AA owns and runs it, only a fifth demonstration scenario is public, and AA's methodology table lists one repeat per task — no averaging over attempts (AA-Briefcase, read 2026-09-08). Its analytical-quality Elo is a preference ranking, which is not a correctness measure.
These are benchmark proxies, not success probabilities. Collaborative problem framing and planning remain unmeasured in Figure; selective escalation and supervision remain unmeasured in Execute. Agent harnesses, tools, task sets and judges differ. Shared AA provenance is not independent corroboration, and small numerical gaps need not matter in real work. A one-benchmark profile carries that scope next to its recommendation.
Data and limitscoverage, effort names, raw rows and sources
Which models are included
The comparison uses two dated captures: the Artificial Analysis model rows of September 6, 2026, which score the Figure lane and the Deliverables focus, and the Coding Agent Index board of September 8, 2026, which scores the Execute lane and Codebase investigation. The broad score table contains exact configurations with explicit lane coverage; the shortlist also includes individual benchmark observations that cannot form a full composite. Each result identifies its measured effort, harness and system; results do not transfer between efforts, harnesses or orthogonal configurations.
An older model stays while something here depends on it: a frontier step, a route, a qualifying shortlist strength or value option, a substitute for a named gap, or an ungraded prediction. Replacement within its vendor line starts that check; age alone never removes it.
Astra without reasoning remains excluded because current OpenAI documentation does not support that configuration. Muse Spark 1.3 max is now publicly available and included; missing costs remain visible. Some Muse native API effort mappings remain unverified, and DeepSeek Flash Vision is experimental. The shortlist states those limits beside each model. The discovery record retains availability checks, omissions and source evidence.
Supporting benchmarks and missing results
Additional evidence can qualify a scoped shortlist alternative without changing the broad score. Figure includes GDP.pdf, CritPt, AnalystAgent and ITBench. Execute includes FrontierCode, Zapier, AA's own Terminal-Bench 2.1 run, SciCode, EnterpriseOps-Gym, AA's partial-credit AutomationBench and APEX-Agents. AA's Terminal-Bench run and the Coding Agent Index's are the same tasks in different harnesses and are kept as separate measurements, never merged. Coverage varies across exact configurations; not measured is a missing result, not poor performance.
AA-Briefcase separately reports analytical quality and presentation. A low presentation score does not arithmetically lower its separate analytical-quality Elo. Its rubric, analytical Elo, and overall Elo answer different questions; none is silently substituted for another here.
Where the numbers come from
Artificial Analysis, its Coding Agent Index, Cognition, and Zapier supply the measurements and accounting. The retained source captures and reproduction commands show how those inputs become the two lane scores, direct costs, and recommendations; the Coding Agent Index capture retains the page as served, its hashes and what it does and does not support. Every page is parsed as data — no downloaded script is executed. Unknown source versions stop regeneration for review. Missing measurements or task accounting remain visible and prevent selection only in the affected lane.
Historical calibrated predictions remain in the dated research and falsifiability records. They do not enter the current comparison or choose a recommendation.
Effort dials, model by model
- GPT-6 Astra has measured
low,medium,high,xhigh, andmaxconfigurations here. OpenAI's model guide excludesnone; Pro is a separate class. - Claude's effort documentation distinguishes model-specific effort settings. The measured Fable configurations include the source's fallback behavior; the direct bill accounts for those fractions.
- Gemini Flash uses
low,medium, andhighthinking levels here. Each is independently measured. - DeepSeek and GLM document
low / high / max; an unmeasured setting stays a named gap. - Qwen3.8-Flash-Next distinguishes its effort dial from disabling thinking. Only the measured
xhighconfiguration is placed here. - Agnes 2.5 Pro Beta uses
reasoninghere to name AA's measured class. That label is not an API effort value.
What the prices do not include
Benchmark API cost is not your invoice. It includes the evaluated runs' recorded API usage, but does not price your extra retries, human review, latency, subscriptions, or the cost of a wrong answer. An API cost ratio is not a subscription-quota ratio. The rail prices a benchmark task mix as shares of a job; it does not forecast the cost of your next job, and the job split it uses is a control you set, not a measurement of you.
| Model · effort | Figure | $ / figure task | Execute | $ / execute task |
|---|---|---|---|---|
| Claude Fable 5.1 · max | 72.68 | $11.9299 | 70.43 | $9.1832 |
Claude Fable 5.1 · xhighMissing evidence
| 71.37 | $9.3321 | Unscored | Unpriced |
| Claude Opus 5 · max | 69.03 | $9.1401 | 67.03 | $8.9435 |
| Claude Opus 5 · xhigh | 68.75 | $7.3447 | 68.15 | $8.1708 |
Claude Fable 5.1 · highMissing evidence
| 68.67 | $6.0158 | Unscored | Unpriced |
Claude Fable 5.1 · mediumMissing evidence
| 66.72 | $4.5761 | Unscored | Unpriced |
| Claude Opus 5 · high | 66.01 | $5.3891 | 65.62 | $3.9186 |
| GPT-6 Astra · max | 65.97 | $4.9964 | 66.97 | $4.7173 |
GPT-6 Astra · xhighMissing evidence
| 64.67 | $3.4940 | Unscored | Unpriced |
| Claude Fable 5 · max | 64.32 | $11.5800 | 67.17 | $11.6874 |
Claude Fable 5.1 · lowMissing evidence
| 63.66 | $3.6276 | Unscored | Unpriced |
GPT-6 Astra · highMissing evidence
| 63.59 | $2.5732 | Unscored | Unpriced |
| Muse Spark 1.3 · max | 62.93 | $1.9245 | 67.97 | $1.5842 |
GPT-6 Astra · mediumMissing evidence
| 62.55 | $2.1481 | Unscored | Unpriced |
GLM-5.3 · maxMissing evidence
| 62.21 | $2.8899 | Unscored | Unpriced |
Kimi K3 · maxMissing evidence
| 61.92 | $3.5065 | Unscored | Unpriced |
| Claude Opus 5 · medium | 61.89 | $2.7798 | 64.07 | $3.1719 |
Grok 4.6 · xhighMissing evidence
| 61.32 | $3.3215 | Unscored | Unpriced |
Grok 4.6 · highMissing evidence
| 60.80 | $2.5600 | Unscored | Unpriced |
| Muse Spark 1.3 · xhigh | 59.25 | $1.7522 | 64.18 | $1.6150 |
Grok 4.6 · mediumMissing evidence
| 58.85 | $1.8538 | Unscored | Unpriced |
| GPT-5.6 Sol · max | 58.50 | $2.1360 | 65.05 | $4.9955 |
GLM-5.3-Flash · maxMissing evidence
| 57.70 | $0.4414 | Unscored | Unpriced |
| GPT-5.6 Sol · xhigh | 56.34 | $1.6340 | 63.34 | $3.7367 |
GPT-6 Astra · lowMissing evidence
| 55.91 | $0.8875 | Unscored | Unpriced |
Qwen3.8 27B · xhighMissing evidence
| 54.08 | $1.1684 | Unscored | Unpriced |
| Muse Spark 1.2 · xhigh | 53.76 | $0.9895 | 61.64 | $2.0728 |
Claude Sonnet 5 · maxMissing evidence
| 53.21 | $7.4509 | Unscored | Unpriced |
| GPT-5.6 Sol · high | 53.20 | $1.1128 | 64.12 | $3.0005 |
| Claude Opus 5 · low | 52.24 | $1.0234 | 59.40 | $2.2952 |
| Claude Opus 4.8 · max | 51.91 | $4.4053 | 62.14 | $7.7247 |
| Grok 4.5 · high | 51.66 | $0.6700 | 64.09 | $2.4390 |
| DeepSeek V4 Pro 0813 · max | 51.21 | $0.5456 | 42.78 | $0.0908 |
Grok 4.6 · lowMissing evidence
| 50.87 | $0.4568 | Unscored | Unpriced |
| GPT-5.6 Terra · max | 50.57 | $1.4991 | 60.42 | $1.9300 |
| GPT-5.6 Terra · xhigh | 49.56 | $0.9580 | 56.03 | $1.3589 |
| GPT-5.6 Luna · max | 49.52 | $0.2077 | 57.17 | $0.2883 |
| GPT-5.5 · xhigh | 49.01 | $2.7488 | 61.03 | $4.7524 |
| GPT-5.6 Sol · medium | 48.61 | $0.5568 | 61.62 | $2.1930 |
| Gemini 3.8 Flash · high | 48.60 | $1.3849 | 61.15 | $2.0378 |
| DeepSeek V4 Flash 0731 · max | 48.46 | $0.2301 | 49.76 | $0.0595 |
GPT-5.5 · highMissing evidence
| 47.30 | $1.7874 | Unscored | Unpriced |
| GPT-5.6 Luna · xhigh | 46.31 | $0.1137 | 52.96 | $0.2359 |
| Gemini 3.7 Flash · high | 45.28 | $1.0991 | 59.63 | $1.2679 |
| GPT-5.6 Terra · high | 44.37 | $0.4746 | 54.65 | $1.1409 |
| Muse Spark 1.1 · xhigh | 43.88 | $1.3011 | 54.92 | $1.4352 |
| GPT-5.5 · medium | 43.23 | $1.0710 | 55.31 | $2.6525 |
| GPT-5.6 Luna · high | 42.45 | $0.0636 | 51.69 | $0.1798 |
| GPT-5.6 Sol · low | 41.36 | $0.2776 | 55.23 | $1.2925 |
Gemini 3.8 Flash · lowMissing evidence
| 40.70 | Unpriced | Unscored | Unpriced |
| GPT-5.6 Terra · medium | 37.70 | $0.2291 | 48.05 | $0.6700 |
| GPT-5.6 Terra · low | 35.00 | $0.1845 | 38.65 | $0.3851 |
Gemini 3.5 Flash · highMissing evidence
| 34.72 | $2.0846 | Unscored | Unpriced |
Gemini 3.5 Flash · mediumMissing evidence
| 34.70 | $2.4744 | Unscored | Unpriced |
| GPT-5.6 Luna · medium | 31.84 | $0.0185 | 41.97 | $0.0884 |
| GPT-5.6 Luna · low | 24.32 | $0.0098 | 25.06 | $0.0389 |
Gemini 3.7 Flash · lowMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Gemini 3.7 Flash · mediumMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Gemini 3.8 Flash · mediumMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Kimi K3 · lowMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Qwen3.8-Flash-Next · xhighMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Agnes 2.5 Pro Beta · reasoningMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
GPT-5.5 · lowMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Qwen3.8 27B · lowMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Qwen3.8 27B · mediumMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
GPT-5.4 Pro · xhighMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
GPT-5.5 Pro · xhighMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Gemini 3.5 Flash · minimalMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
DeepSeek V4 Flash Vision · maxMissing evidence
| Unscored | Unpriced | Unscored | Unpriced |
Claude Fable 5 · highMissing evidence
| Unscored | Unpriced | 65.13 | $5.9710 |
Claude Fable 5 · lowMissing evidence
| Unscored | Unpriced | 58.56 | $3.1661 |
Claude Fable 5 · mediumMissing evidence
| Unscored | Unpriced | 63.20 | $4.7355 |
Claude Fable 5 · xhighMissing evidence
| Unscored | Unpriced | 66.02 | $8.5258 |
Kimi K3 · unspecifiedMissing evidence
| Unscored | Unpriced | 62.64 | $3.0816 |
GPT-5.6 Luna · noneMissing evidence
| Unscored | Unpriced | 19.10 | $0.0695 |
Claude Opus 4.8 · highMissing evidence
| Unscored | Unpriced | 57.58 | $3.7823 |
Claude Opus 4.8 · lowMissing evidence
| Unscored | Unpriced | 49.04 | $2.1826 |
Claude Opus 4.8 · mediumMissing evidence
| Unscored | Unpriced | 55.77 | $3.3038 |
Claude Opus 4.8 · xhighMissing evidence
| Unscored | Unpriced | 59.07 | $5.6654 |
Claude Opus 5 · noneMissing evidence
| Unscored | Unpriced | 59.17 | $3.5309 |
Qwen3.8 Max · unspecifiedMissing evidence
| Unscored | Unpriced | 61.31 | $3.2306 |
GPT-5.6 Sol · noneMissing evidence
| Unscored | Unpriced | 43.37 | $1.0869 |
GPT-5.6 Terra · noneMissing evidence
| Unscored | Unpriced | 23.09 | $0.2927 |