LLM landscape
Model benchmark performance
How models score on individual public evaluations. Leaderboards are curated from published scores; a scheduled refresh republishes the current table but does not fill evaluations whose scores have not been added.
- Data as of
- Aug 25, 2026
10
1
13
Coverage is partial: 1 of 10 tracked evaluations has scores. 9 await curated leaderboard data. A refresh republishes the curated snapshot; it does not generate missing scores.
Evaluation
GDPval (Artificial Analysis)
Performance on economically valuable, expert-graded real-world work tasks (Artificial Analysis' GDPval run); score out of 3000. · higher is better
On GDPval (Artificial Analysis), Claude Fable 5 leads with 1,932 (Anthropic). Anthropic holds 4 of the top 5.
13 models scored
GDPval (Artificial Analysis): Performance on economically valuable, expert-graded real-world work tasks (Artificial Analysis' GDPval run); score out of 3000. Each evaluation is a published Artificial Analysis benchmark; one bar per model, height = the model's reported score on the selected evaluation (higher is better). Scores are a curated, source-linked table copied verbatim from the public leaderboards — no estimates. An evaluation with no rows yet is shown as awaiting data rather than populated with guesses.