LLM landscape

Compute vs capability

Track how published training compute relates to model capability in one dataset. Each point carries a status and source status so measured rows and explicit assumptions stay inspectable without splitting the view.

Data as of
Aug 25, 2026
Measured

42

Assumptions

15

Review queue

84

Compute vs capability

Training compute vs capability

Aug 25, 2026, 5:20 AM

57 models place their published training compute against capability. Across 42 measured models, capability climbs about +9.5 ECI points per 10× of training compute. GPT-5.5 Pro sits at the capability frontier (162 ECI).

View and filters

Each dot is a model · log compute axis · higher and right-er means more capability per unit of training compute

901001101201301401501601701023102410251026Training compute, FLOP (log scale)Capability level (ECI score)Falcon-180BYi-34BMixtral 8x7BAmazon Nova ProPhi-4Llama 4 MaverickGLM-4.6Nemotron 3 UltraQwen3-235B-A22B-Thinking (Jul 2025)GLM-5Grok 4InklingDeepSeek-V4-ProGemini 3.1 ProKimi K3Claude Opus 4.8GPT-5.5 Pro
Organisation
MarkerOpen weightsProprietaryAssumption (outlined)

Training compute is Epoch's published FLOP estimate on a log axis; capability is the ECI score. Points are coloured by organisation and shaped by weights access (square = open weights, circle = proprietary); outlined markers carry an assumed rather than measured value. The dashed trend line is a log-linear fit over visible measured rows only, so assumption rows plot but do not bend it. Hover near a point to inspect it, click to pin, Escape to clear, use legend filters, and Frontier focus to zoom crowded high-capability clusters.