LLM landscape

Model benchmark performance

How models score on individual public evaluations. Leaderboards are curated from published scores; a scheduled refresh republishes the current table but does not fill evaluations whose scores have not been added.

Data as of
Aug 25, 2026
Evaluations

10

Populated

1

Models

13

Coverage is partial: 1 of 10 tracked evaluations has scores. 9 await curated leaderboard data. A refresh republishes the curated snapshot; it does not generate missing scores.

Evaluation

GDPval (Artificial Analysis)

Performance on economically valuable, expert-graded real-world work tasks (Artificial Analysis' GDPval run); score out of 3000. · higher is better

On GDPval (Artificial Analysis), Claude Fable 5 leads with 1,932 (Anthropic). Anthropic holds 4 of the top 5.

13 models scored

1,932Claude Fable 5
1,890Claude Opus 4.8
1,656Gemini 3.5 Flash
1,633Claude Sonnet 4.6
1,606Claude Opus 4.6
1,581MiMo-V2.5-Pro
1,554DeepSeek-V4-Pro-Max
1,494MiniMax M2.7
1,444Muse Spark
1,426MiMo-V2-Pro
1,410MiMo-V2-Omni
1,395DeepSeek-V4-Flash-Max
1,317Gemini 3.1 Pro
OrganisationAnthropicGoogleXiaomiDeepSeekMiniMaxMeta

GDPval (Artificial Analysis): Performance on economically valuable, expert-graded real-world work tasks (Artificial Analysis' GDPval run); score out of 3000. Each evaluation is a published Artificial Analysis benchmark; one bar per model, height = the model's reported score on the selected evaluation (higher is better). Scores are a curated, source-linked table copied verbatim from the public leaderboards — no estimates. An evaluation with no rows yet is shown as awaiting data rather than populated with guesses.