Skip to content

AI Model Evaluation

LLM Model Benchmarks

Current coding leaders

Rank 181.4

Claude Fable 5

$2.5/task

Rank 279

GPT-5.6 Sol

$1.4/task

Rank 378.6

Kimi K3

$0.75/task

A work-first guide to professional, agentic, coding, regulated-domain, STEM, preference, safety, and cost benchmarks. Current score cards use the BenchLM public feed; Artificial Analysis is shown separately as an independent evaluation suite.

Updated 2026-08-30

Benchmark map

All benchmark types

A compact map of the evals worth knowing before comparing Claude, ChatGPT, Gemini, open-weight models, or any new frontier release.

Professional work

Business, legal, finance, healthcare, research, documents, and human-facing output.

Current best

80.1Claude Opus 5

$1.25/task

#2 75.5 · $2.5/task · Claude Fable 5

#3 75.2 · $0.3/task · Qwen3.8 Max

Design

Design fidelity, human preference, working web applications, and presentation quality.

Current best

38.3%Claude Opus 4.6

FigmaBench

#2 23.3% · Mistral Large 3

#3 19.9% · Kimi K2.5

Security

Offensive and defensive capability, vulnerability work, secure coding, and guarded deployment.

Current best

95.5Qwen3.8 Max

$0.3/task

#2 85.2 · $1.4/task · GPT-5.6 Sol

#3 83.5 · $1.25/task · Claude Opus 5

Coding

Repository repair, fresh code generation, and terminal execution.

Current best

81.4Claude Fable 5

$2.5/task

#2 79 · $1.4/task · GPT-5.6 Sol

#3 78.6 · $0.75/task · Kimi K3

Connectivity

Tools, browsers, interfaces, retrieval systems, and sustained context.

Current best

80.1Claude Opus 5

$1.25/task

#2 75.5 · $2.5/task · Claude Fable 5

#3 75.2 · $0.3/task · Qwen3.8 Max

STEM

Broad intelligence, science, mathematics, and difficult held-out reasoning tasks.

Current best

95.5Qwen3.8 Max

$0.3/task

#2 85.2 · $1.4/task · GPT-5.6 Sol

#3 83.5 · $1.25/task · Claude Opus 5

Alignment

Whether a model follows intent or cheats, deceives, schemes, sandbags, or games an evaluation.

Current best

83.2Claude Fable 5

$2.5/task

#2 83.1 · $1.25/task · Claude Opus 5

#3 82.2 · $1.4/task · GPT-5.6 Sol

Cost

Cost per completed task, speed, retries, reliability, and operational efficiency.

Current best

100Trinity-Large-Preview

$0.06/task

#2 79 · $0.07/task · DeepSeek V4 Pro 0813

#3 52.4 · $0.11/task · Gemini 3.5 Flash-Lite

Selector

Pick the work, then pick the benchmarks

Choose the outcomes that matter. The selector returns the benchmark families that should carry the most weight.

Selected work winner

Claude Opus 5

$1.25/task estimated

Fit score

80.2

Based on the selected work. Cost changes the ranking only when Cost Per Task is selected.

Choose priorities

Relevance to task

Terminal and developer-environment agents

Function calling and tool selection

Fresh coding and contest programming

Repository issue repair

Long-context retrieval and memory

Professional workflow execution

Human preference arenas

Model snapshot

Current BenchLM scores

BenchLM snapshot fetched 2026-08-30. Restricted-access and unknown-availability models are excluded; the comparison keeps overall leaders plus top public Anthropic, OpenAI, Google/Gemini, and DeepSeek coverage.

RankModelOverall
1Claude Fable 5Anthropic1M / $10/$5083.2
2Claude Opus 5Anthropic1M / $5/$2583.1
3GPT-5.6 SolOpenAI1.05M / $5/$3082.2
4Kimi K3Moonshot AI1.05M / $3/$1580.6
5Qwen3.8 MaxAlibaba1M / $1.6/$4.879.2
6Hy4 previewTencent1M / $0/$079.2
7Muse Spark 1.1Meta1M / $1.25/$4.2577.1
8Claude Opus 4.8Anthropic1M / $5/$2576.6
9Gemini 3.6 FlashGoogle1M / $1.5/$7.575.7
10Grok 4.5xAI500K / $2/$675.7
11GPT-5.4OpenAI1.05M / $2.5/$1573.6
12GPT-5.6 TerraOpenAI1.05M / $2.5/$1573
13dots3-note PreviewDots Studio512K / $0/$068.7
14Gemini 3 ProGoogle2M / $2/$1267.6
15Gemini 3.5 Flash-LiteGoogle1M / $0.3/$2.565.5
16DeepSeek V4 Pro 0813DeepSeek1M / $0.44/$0.8761
17Qwen 3.6 Max (preview)Alibaba246K / $1.04/$6.2460.4
18Trinity-Large-PreviewArcee AI131K / $0.25/$156.7
19Hy3 PreviewTencent256K / $0/$043.7

FAQ

Quick answers

Claude Fable 5 vs GPT-5.6 Sol vs Kimi K3

Claude Fable 5 currently has the highest live software-engineering score shown here: 81.4. Restricted-access and unknown-availability models are excluded. Use it as the first model to inspect for coding work, then validate with repo repair, terminal-agent, and local codebase tasks.

What is the most trustworthy LLM benchmark?

There is no single most trustworthy benchmark. The best benchmark is the one that closely matches the task, uses a transparent and consistent harness, is fresh enough to reduce contamination, and reports enough metadata to compare models fairly.

Should I trust vendor-published benchmark results?

Trust them as useful claims, not final evidence. Provider system cards and launch posts often have the freshest model data, but the result should be checked against benchmark-owner leaderboards, third-party runs, and local evaluations before it drives a production decision.

Why do two pages show different scores for the same model and benchmark?

Usually because the benchmark version, subset, scaffold, tools, effort level, context budget, number of attempts, or scoring rule changed. Same benchmark name does not guarantee same experiment.

What benchmark matters most for coding agents?

Use several: audited repository-repair tasks, Terminal-Bench for shell and tool persistence, LiveCodeBench for fresh coding ability, and local repo tasks for the work that matters to your team. Treat SWE-Bench Pro cautiously while its reported task-quality problems remain unresolved.

How should token cost be evaluated?

Use cost per completed task, not price per million tokens alone. Count input, cached input, output, reasoning tokens when billed, tool costs, retries, long-context surcharges, latency, and the cost of human correction.

Inspect first

Sources

How to read the evidence

Benchmark owners

Trust the exact task they run; never treat one task as a universal rank.

AA + Vals

Use for comparable agent systems and finished professional work.

METR

Use for long-horizon autonomy, monitorability, and frontier risk.

LMArena

Use for human preference and feel, not correctness.

BenchLM

Use as broad, fresh radar; inspect the underlying evidence before deciding.

Third-party data note: live rows come from public benchmark and pricing feeds, not internal Dreamers testing. Restricted-access and unknown-availability rows are excluded from public leaders; public previews remain public. Cost per task estimates use 100K input and 30K output tokens; tools, retries, cache, and failed attempts are excluded.