Claude
Claude Fable 5.1
Public rank 1 · Proprietary
Overall 84.4
Coding 84.2
Agentic 80.1
Reasoning 79.4
AI Model Comparison
Latest public provider picks: Claude is Claude Fable 5.1, ChatGPT is GPT-6 Astra, and Gemini is Gemini 3.8 Flash.
A current public benchmark comparison of latest Claude, ChatGPT/GPT, and Gemini models, with what the scores mean for coding, automation, reasoning, and cost per task.
Updated 2026-09-09
Current comparison
The table uses public benchmark rows and excludes restricted-access and unknown-availability models before selecting each provider.
Claude
Claude Fable 5.1
Public rank 1 · Proprietary
Overall 84.4
Coding 84.2
Agentic 80.1
Reasoning 79.4
ChatGPT
GPT-6 Astra
Public rank 2 · Proprietary
Overall 84.1
Coding 74.6
Agentic 70.4
Reasoning 89.5
Gemini
Gemini 3.8 Flash
Public rank 6 · Proprietary
Overall 77.8
Coding 67.7
Agentic 66.5
Reasoning 76.9
How to read it
Overall score is a shortlist signal. Coding, agentic workflow, reasoning, context, and token price decide whether the right model changes by task.
A high coding score is the first filter, not the final decision. Repo repair, terminal-agent behavior, long context, and local tests decide whether the model can actually move a codebase.
Workflow automation rewards tool selection, persistence, browser or terminal actions, and recovery from bad intermediate results. That can favor a different model than raw coding or reasoning.
Token price matters only after accounting for retries, long context, output length, cached input, tool calls, and human correction. Cheap tokens can be expensive if the task fails.
Method
For production, choose the model by task: repo repair, automation, retrieval, regulated-domain work, latency, and cost per completed task.
This page is intentionally narrow. It compares the latest public Claude, ChatGPT/GPT, and Gemini rows, then points back to the full benchmark guide for benchmark-family selection, cost interpretation, and local evaluation design.
Read the benchmark guide ->Inspect first
Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.