Skip to content

AI Model Watch

All About Claude Fable 5.1

Current public winner: Claude Fable 5.1 with an overall score of 83. Public #2: GPT-6 Astra.

Public models only

Claude Fable 5.1 is the current public frontier-model leader in the live benchmark snapshot. Restricted-access rows move to the unreleased watch page and unknown availability is withheld, so this dossier can show the public #1, public #2, category scores, costs, and nearest challengers. Updated 2026-09-04.

Updated 2026-09-04

Public #1

83

Public rank 1 in the current snapshot.

Public #2

GPT-6 Astra

Overall 81.1 · rank 2.

Coding

84.2

Tied #1 on the coding signal.

Cost per task

$2.5

Estimate for 100K input + 30K output tokens; tools, retries, and cache excluded.

Daily verdict

Claude Fable 5.1 is the public model to inspect first

That does not make it the best model for every job. It means the public side of the benchmark feed currently says it deserves the first serious look.

What the scoreboard is really saying

Claude Fable 5.1 leads this live public comparison with an overall score of 83 and a coding score of 84.2, plus an agentic score of 78.7.

Read that as a selection signal, not a guarantee. The model still needs to be tested on your repo, your tools, your risk posture, and your budget. The public score is the shortlist; the task result is the decision.

Public #2 is GPT-6 Astra with an overall score of 81.1.

Restricted-access rows are handled separately so they do not distort the public-model answer. View the unreleased watch page.

There is a real tie signal in the current data. If another model shares the top category score, the tiebreak comes from public rank, overall score, and how much the category actually matters for the work.

Model card

Creator

Anthropic

Access

Public

Context

1M

Reasoning

#3 · 79.4

Input $/M

$10

Output $/M

$50

Scoreboard

Where it wins, ties, and gets challenged

A useful model page should show the public model and the field around it. The nearest rivals matter because ties and narrow gaps are where local testing changes the answer.

Overall leaders

#1

Claude Fable 5.1

Anthropic · dossier model

83

#2

GPT-6 Astra

OpenAI

81.1

#3

Claude Fable 5

Anthropic

80.9

#4

Claude Opus 5

Anthropic

80.7

#5

GPT-5.6 Sol

OpenAI

79.7

Coding leaders

#1

Claude Fable 5.1

Anthropic · dossier model

84.2

#2

Claude Fable 5

Anthropic

76.9

#3

Claude Opus 5

Anthropic

75.6

#4

GPT-6 Astra

OpenAI

75.3

#5

GPT-5.6 Sol

OpenAI

74.4

Agentic leaders

#1

Claude Fable 5.1

Anthropic · dossier model

78.7

#2

Claude Opus 5

Anthropic

77.4

#3

Claude Fable 5

Anthropic

74.8

#4

Kimi K3

Moonshot AI

71.8

#5

GPT-6 Astra

OpenAI

70.3

Reasoning leaders

#1

GPT-6 Astra

OpenAI

88.8

#2

Qwen3.8 Max

Alibaba

86.5

#3

Claude Fable 5.1

Anthropic · dossier model

79.4

#4

Gemini 3.8 Flash

Google

78.5

#5

Kimi K3

Moonshot AI

78.5

The field

Closest public rivals

These are the nearby public rows after restricted-access and unknown-availability models are withheld.

GPT-6 Astra

OpenAI · public rank 2

Overall 81.1

Coding 75.3

Agentic 70.3

$/M $10 / $50

Claude Fable 5

Anthropic · public rank 3

Overall 80.9

Coding 76.9

Agentic 74.8

$/M $10 / $50

Claude Opus 5

Anthropic · public rank 4

Overall 80.7

Coding 75.6

Agentic 77.4

$/M $5 / $25

GPT-5.6 Sol

OpenAI · public rank 5

Overall 79.7

Coding 74.4

Agentic 70.1

$/M $5 / $30

Gemini 3.8 Flash

Google · public rank 6

Overall 78.4

Coding 68.1

Agentic 67.6

$/M $0.75 / $3.75

Use it well

How to turn a winner into a decision

The live winner is the beginning of evaluation, not the end.

Use the public rank to shortlist

Start with the public leader and closest rivals. Do not spend local eval time on every model when the public field already narrows the search.

Use local tests to decide

Run the winner against your repo, RAG corpus, browser workflow, security posture, and cost envelope. Public benchmarks cannot see your failure modes.

Use cost per completed task

Token price is only useful after retries, output length, long context, tool calls, latency, and human correction are counted.

FAQ

Quick answers

What does best frontier model today mean?

It means the current benchmark leader among models with evidence of broad public access. Restricted and unknown-availability rows are excluded. It is a daily shortlist signal, not a permanent recommendation.

Where do preview models go?

A public preview stays on this page; the label alone does not prove restricted access. Models with verified partner-only or private access go to the unreleased watch page. Unknown availability is withheld.

Can the top model change without this page being redeployed?

Yes when the route is served through the app-server SSR path or when the browser hydrates from the live public feeds. Static source HTML remains the fallback until this page is moved behind the SSR origin.

Why can one model be best overall while another is best for coding?

Overall score averages broad public signals. Coding, agentic workflow, reasoning, context, and cost can move independently, so a model can be the overall leader while a different model is the better choice for a narrow job.

Should cost decide the winner?

Only for cost-sensitive workloads. Token price matters after accounting for context length, output length, retries, tool calls, cached input, latency, and human correction. The cheapest model can become expensive if it fails the task.

Inspect first

Sources

Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.