Claude Fable 5
$2.5/task
AI Model Evaluation
Current coding leaders
$2.5/task
$1.4/task
$0.75/task
A work-first guide to professional, agentic, coding, regulated-domain, STEM, preference, safety, and cost benchmarks. Current score cards use the BenchLM public feed; Artificial Analysis is shown separately as an independent evaluation suite.
Updated 2026-08-30
Benchmark map
A compact map of the evals worth knowing before comparing Claude, ChatGPT, Gemini, open-weight models, or any new frontier release.
Professional work
Business, legal, finance, healthcare, research, documents, and human-facing output.
Current best
$1.25/task
#2 75.5 · $2.5/task · Claude Fable 5
#3 75.2 · $0.3/task · Qwen3.8 Max
Design
Design fidelity, human preference, working web applications, and presentation quality.
Current best
FigmaBench
#2 23.3% · Mistral Large 3
#3 19.9% · Kimi K2.5
Security
Offensive and defensive capability, vulnerability work, secure coding, and guarded deployment.
Current best
$0.3/task
#2 85.2 · $1.4/task · GPT-5.6 Sol
#3 83.5 · $1.25/task · Claude Opus 5
Coding
Repository repair, fresh code generation, and terminal execution.
Current best
$2.5/task
#2 79 · $1.4/task · GPT-5.6 Sol
#3 78.6 · $0.75/task · Kimi K3
Connectivity
Tools, browsers, interfaces, retrieval systems, and sustained context.
Current best
$1.25/task
#2 75.5 · $2.5/task · Claude Fable 5
#3 75.2 · $0.3/task · Qwen3.8 Max
STEM
Broad intelligence, science, mathematics, and difficult held-out reasoning tasks.
Current best
$0.3/task
#2 85.2 · $1.4/task · GPT-5.6 Sol
#3 83.5 · $1.25/task · Claude Opus 5
Alignment
Whether a model follows intent or cheats, deceives, schemes, sandbags, or games an evaluation.
Current best
$2.5/task
#2 83.1 · $1.25/task · Claude Opus 5
#3 82.2 · $1.4/task · GPT-5.6 Sol
Cost
Cost per completed task, speed, retries, reliability, and operational efficiency.
Current best
$0.06/task
#2 79 · $0.07/task · DeepSeek V4 Pro 0813
#3 52.4 · $0.11/task · Gemini 3.5 Flash-Lite
Selector
Choose the outcomes that matter. The selector returns the benchmark families that should carry the most weight.
Selected work winner
Claude Opus 5
$1.25/task estimated
Fit score
80.2
Based on the selected work. Cost changes the ranking only when Cost Per Task is selected.
Choose priorities
Relevance to task
Model snapshot
BenchLM snapshot fetched 2026-08-30. Restricted-access and unknown-availability models are excluded; the comparison keeps overall leaders plus top public Anthropic, OpenAI, Google/Gemini, and DeepSeek coverage.
| Rank | Model | Overall | Coding | Agentic | Reasoning | Context | $/M In/Out |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic1M / $10/$50 | 83.2 | 81.4 | 75.5 | NA | 1M | $10 / $50 |
| 2 | Claude Opus 5Anthropic1M / $5/$25 | 83.1 | 78.4 | 80.1 | 83.5 | 1M | $5 / $25 |
| 3 | GPT-5.6 SolOpenAI1.05M / $5/$30 | 82.2 | 79 | 68.9 | 85.2 | 1.05M | $5 / $30 |
| 4 | Kimi K3Moonshot AI1.05M / $3/$15 | 80.6 | 78.6 | 73.6 | NA | 1.05M | $3 / $15 |
| 5 | Qwen3.8 MaxAlibaba1M / $1.6/$4.8 | 79.2 | 65.8 | 75.2 | 95.5 | 1M | $1.6 / $4.8 |
| 6 | Hy4 previewTencent1M / $0/$0 | 79.2 | 68.9 | 70 | NA | 1M | $0 / $0 |
| 7 | Muse Spark 1.1Meta1M / $1.25/$4.25 | 77.1 | 65.7 | 59.8 | NA | 1M | $1.25 / $4.25 |
| 8 | Claude Opus 4.8Anthropic1M / $5/$25 | 76.6 | 72 | 61.8 | 68.3 | 1M | $5 / $25 |
| 9 | Gemini 3.6 FlashGoogle1M / $1.5/$7.5 | 75.7 | 65.3 | 38.3 | NA | 1M | $1.5 / $7.5 |
| 10 | Grok 4.5xAI500K / $2/$6 | 75.7 | 57.9 | 63.1 | 52.1 | 500K | $2 / $6 |
| 11 | GPT-5.4OpenAI1.05M / $2.5/$15 | 73.6 | 58.9 | 57.5 | 69.8 | 1.05M | $2.5 / $15 |
| 12 | GPT-5.6 TerraOpenAI1.05M / $2.5/$15 | 73 | 69.3 | 58.9 | 78.1 | 1.05M | $2.5 / $15 |
| 13 | dots3-note PreviewDots Studio512K / $0/$0 | 68.7 | 61.5 | 61.1 | 76 | 512K | $0 / $0 |
| 14 | Gemini 3 ProGoogle2M / $2/$12 | 67.6 | 62 | 59.1 | 34.2 | 2M | $2 / $12 |
| 15 | Gemini 3.5 Flash-LiteGoogle1M / $0.3/$2.5 | 65.5 | 52.7 | 31 | 54.3 | 1M | $0.3 / $2.5 |
| 16 | DeepSeek V4 Pro 0813DeepSeek1M / $0.44/$0.87 | 61 | 50.5 | 59 | NA | 1M | $0.44 / $0.87 |
| 17 | Qwen 3.6 Max (preview)Alibaba246K / $1.04/$6.24 | 60.4 | 58 | 58.2 | NA | 246K | $1.04 / $6.24 |
| 18 | Trinity-Large-PreviewArcee AI131K / $0.25/$1 | 56.7 | NA | NA | NA | 131K | $0.25 / $1 |
| 19 | Hy3 PreviewTencent256K / $0/$0 | 43.7 | 47.1 | 47.1 | NA | 256K | $0 / $0 |
FAQ
Claude Fable 5 currently has the highest live software-engineering score shown here: 81.4. Restricted-access and unknown-availability models are excluded. Use it as the first model to inspect for coding work, then validate with repo repair, terminal-agent, and local codebase tasks.
There is no single most trustworthy benchmark. The best benchmark is the one that closely matches the task, uses a transparent and consistent harness, is fresh enough to reduce contamination, and reports enough metadata to compare models fairly.
Trust them as useful claims, not final evidence. Provider system cards and launch posts often have the freshest model data, but the result should be checked against benchmark-owner leaderboards, third-party runs, and local evaluations before it drives a production decision.
Usually because the benchmark version, subset, scaffold, tools, effort level, context budget, number of attempts, or scoring rule changed. Same benchmark name does not guarantee same experiment.
Use several: audited repository-repair tasks, Terminal-Bench for shell and tool persistence, LiveCodeBench for fresh coding ability, and local repo tasks for the work that matters to your team. Treat SWE-Bench Pro cautiously while its reported task-quality problems remain unresolved.
Use cost per completed task, not price per million tokens alone. Count input, cached input, output, reasoning tokens when billed, tool costs, retries, long-context surcharges, latency, and the cost of human correction.
Inspect first
How to read the evidence
Benchmark owners
Trust the exact task they run; never treat one task as a universal rank.
AA + Vals
Use for comparable agent systems and finished professional work.
METR
Use for long-horizon autonomy, monitorability, and frontier risk.
LMArena
Use for human preference and feel, not correctness.
BenchLM
Use as broad, fresh radar; inspect the underlying evidence before deciding.
Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.