GPT-6 Astra
OpenAI · public rank 2
Overall 81.1
Coding 75.3
Agentic 70.3
$/M $10 / $50
AI Model Watch
Current public winner: Claude Fable 5.1 with an overall score of 83. Public #2: GPT-6 Astra.
Public models only
Claude Fable 5.1 is the current public frontier-model leader in the live benchmark snapshot. Restricted-access rows move to the unreleased watch page and unknown availability is withheld, so this dossier can show the public #1, public #2, category scores, costs, and nearest challengers. Updated 2026-09-04.
Updated 2026-09-04
Public #1
83
Public rank 1 in the current snapshot.
Public #2
GPT-6 Astra
Overall 81.1 · rank 2.
Coding
84.2
Tied #1 on the coding signal.
Cost per task
$2.5
Estimate for 100K input + 30K output tokens; tools, retries, and cache excluded.
Daily verdict
That does not make it the best model for every job. It means the public side of the benchmark feed currently says it deserves the first serious look.
Claude Fable 5.1 leads this live public comparison with an overall score of 83 and a coding score of 84.2, plus an agentic score of 78.7.
Read that as a selection signal, not a guarantee. The model still needs to be tested on your repo, your tools, your risk posture, and your budget. The public score is the shortlist; the task result is the decision.
Public #2 is GPT-6 Astra with an overall score of 81.1.
Restricted-access rows are handled separately so they do not distort the public-model answer. View the unreleased watch page.
There is a real tie signal in the current data. If another model shares the top category score, the tiebreak comes from public rank, overall score, and how much the category actually matters for the work.
Model card
Creator
Anthropic
Access
Public
Context
1M
Reasoning
#3 · 79.4
Input $/M
$10
Output $/M
$50
Scoreboard
A useful model page should show the public model and the field around it. The nearest rivals matter because ties and narrow gaps are where local testing changes the answer.
Overall leaders
#1
Claude Fable 5.1
Anthropic · dossier model
83
#2
GPT-6 Astra
OpenAI
81.1
#3
Claude Fable 5
Anthropic
80.9
#4
Claude Opus 5
Anthropic
80.7
#5
GPT-5.6 Sol
OpenAI
79.7
Coding leaders
#1
Claude Fable 5.1
Anthropic · dossier model
84.2
#2
Claude Fable 5
Anthropic
76.9
#3
Claude Opus 5
Anthropic
75.6
#4
GPT-6 Astra
OpenAI
75.3
#5
GPT-5.6 Sol
OpenAI
74.4
Agentic leaders
#1
Claude Fable 5.1
Anthropic · dossier model
78.7
#2
Claude Opus 5
Anthropic
77.4
#3
Claude Fable 5
Anthropic
74.8
#4
Kimi K3
Moonshot AI
71.8
#5
GPT-6 Astra
OpenAI
70.3
Reasoning leaders
#1
GPT-6 Astra
OpenAI
88.8
#2
Qwen3.8 Max
Alibaba
86.5
#3
Claude Fable 5.1
Anthropic · dossier model
79.4
#4
Gemini 3.8 Flash
78.5
#5
Kimi K3
Moonshot AI
78.5
The field
These are the nearby public rows after restricted-access and unknown-availability models are withheld.
GPT-6 Astra
OpenAI · public rank 2
Overall 81.1
Coding 75.3
Agentic 70.3
$/M $10 / $50
Claude Fable 5
Anthropic · public rank 3
Overall 80.9
Coding 76.9
Agentic 74.8
$/M $10 / $50
Claude Opus 5
Anthropic · public rank 4
Overall 80.7
Coding 75.6
Agentic 77.4
$/M $5 / $25
GPT-5.6 Sol
OpenAI · public rank 5
Overall 79.7
Coding 74.4
Agentic 70.1
$/M $5 / $30
Gemini 3.8 Flash
Google · public rank 6
Overall 78.4
Coding 68.1
Agentic 67.6
$/M $0.75 / $3.75
Use it well
The live winner is the beginning of evaluation, not the end.
Start with the public leader and closest rivals. Do not spend local eval time on every model when the public field already narrows the search.
Run the winner against your repo, RAG corpus, browser workflow, security posture, and cost envelope. Public benchmarks cannot see your failure modes.
Token price is only useful after retries, output length, long context, tool calls, latency, and human correction are counted.
FAQ
It means the current benchmark leader among models with evidence of broad public access. Restricted and unknown-availability rows are excluded. It is a daily shortlist signal, not a permanent recommendation.
A public preview stays on this page; the label alone does not prove restricted access. Models with verified partner-only or private access go to the unreleased watch page. Unknown availability is withheld.
Yes when the route is served through the app-server SSR path or when the browser hydrates from the live public feeds. Static source HTML remains the fallback until this page is moved behind the SSR origin.
Overall score averages broad public signals. Coding, agentic workflow, reasoning, context, and cost can move independently, so a model can be the overall leader while a different model is the better choice for a narrow job.
Only for cost-sensitive workloads. Token price matters after accounting for context length, output length, retries, tool calls, cached input, latency, and human correction. The cheapest model can become expensive if it fails the task.
Inspect first
Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.