AI Model Watch
All About Claude Mythos 5
Last qualifying restricted challenger: Claude Mythos 5 scored 83.4 against Claude Fable 5 at 83.1.
Restricted-access watch
Claude Mythos 5 is the last restricted-access model recorded beating the public frontier. Its qualifying score is preserved until another restricted model clears the current public leader. Updated 2026-09-04.
Updated 2026-09-04
Restricted #1
83.4
Cleared Claude Fable 5 at 83.1 on 2026-09-01.
Restricted #2
NA
No second row in this snapshot.
Coding
81.4
Tied #1 on the coding signal.
Cost per task
$2.5
Estimate for 100K input + 30K output tokens; tools, retries, and cache excluded.
Daily verdict
Claude Mythos 5 remains the last challenger to clear the public frontier
No currently restricted row cleared the public leader, so the last qualifying result is preserved with its original comparison date.
What the scoreboard is really saying
Claude Mythos 5 last qualified restricted-access comparison with an overall score of 83.4 and a coding score of 81.4, plus an agentic score of 75.1.
This score is preserved from 2026-09-01. The current feed was checked again, but no restricted model beat today’s public leader.
For production-facing model selection, compare this watchlist against the public winner. View the public model page.
There is a real tie signal in the current data. If another model shares the top category score, the tiebreak comes from public rank, overall score, and how much the category actually matters for the work.
Model card
Creator
Anthropic
Access
Restricted / partner only
Context
1M
Reasoning
NA
Input $/M
$10
Output $/M
$50
Scoreboard
Where it wins, ties, and gets challenged
This page only admits verified restricted-access models that beat the public leader. When none do, it preserves the last qualifying challenger rather than promoting a weaker row.
Overall leaders
#1
Claude Mythos 5
Anthropic · dossier model
83.4
Coding leaders
#1
Claude Mythos 5
Anthropic · dossier model
81.4
Agentic leaders
#1
Claude Mythos 5
Anthropic · dossier model
75.1
The field
Closest restricted-access rivals
Only verified restricted-access rows that also clear the public leader appear here. Unknown and weaker restricted rows are withheld.
No second restricted-access row in this snapshot.
The section will populate automatically when the source feed exposes another comparable restricted-access model that beats the public leader.
Use it well
How to turn a winner into a decision
The live winner is the beginning of evaluation, not the end.
Use restricted models as horizon signals
A restricted leader can show where the market is moving, but it should not replace public-model evaluation until access and terms are real.
Compare against public #1 and #2
A restricted score only matters if it changes the decision against the best broadly available models.
Wait for broad access
Availability, pricing, rate limits, safety behavior, and model identity can change before restricted access becomes broadly public.
FAQ
Quick answers
What counts as unreleased here?
The stable URL uses “unreleased,” but the test is narrower: the model must have primary-source evidence of restricted, partner-only, private, or otherwise non-broad access, and its overall score must beat the current public leader. Preview and beta names alone do not qualify.
What is dots3-note Preview, and why is it not here?
It is a public open-weight multimodal model from Dots Studio. Its official model card publishes downloadable weights under Apache 2.0, so “Preview” is a version label, not evidence that the model is unreleased.
Can an unreleased model be the real best model?
It can be the highest-scoring row in the feed, but that is different from being the best model most teams can buy or deploy. Access, terms, limits, and pricing can change before broad availability.
What happens when no restricted model beats the public leader?
The page keeps the last restricted model that qualified, preserving its score, the public model it beat, and the original comparison date. The latest feed-check date remains separate so old evidence is not presented as a new result.
Why keep this separate from best frontier model today?
The main page answers what broadly available model to inspect first. This page answers what restricted-access model is setting the visible frontier.
Should teams test restricted-access models?
Yes, when they have legitimate access and a clear reason. Test restricted models against the same local tasks as public models, but do not build a production dependency before access and terms are durable.
Inspect first
Sources
- BenchLM public leaderboard endpoint
- BenchLM public pricing endpoint
- Models.dev model database
- Anthropic Claude Mythos 5 availability
- dots3-note Preview public open-weight model card
- Kimi K3 public release
- Best public frontier model page
- LLM model benchmarks guide
- LMArena Leaderboard
- SWE-bench repository
- Terminal-Bench 2.1
Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.