Open-model infrastructure

LLM hardware requirements

Compare the hardware needed for downloadable models across precision, quantization method, runtime, workload profile, and source strength. Each rank represents one concrete configuration, not an absolute barrier.

Data as of
Aug 25, 2026
Models

16

Configurations

44

Verified recipes

40

Quality comparisons

0

Hardware ladder

Model configurations by deployment footprint

Counts overlap when a model has several validated or calculated configurations.

Local

CPU / unified

6 models

1 GPU

Single accelerator

7 models

2–4 GPUs

Multi-GPU

5 models

GPU server

5–8 GPUs

9 models

Cluster

9+ GPUs

1 model

Weight precision

Showing published configurations.

Ranked configurations

16 models match

One primary configuration per model · alternatives expand in place

RankModelVariantHardwareProfileQuality deltaEvidence
1

Gemma 4 E2B

Google

5B total · 5B active

INT4INT4

vLLM

75% less weight memory

4 GB unified memory

4 GB memory floor

Context not stated

loadable
No comparable result published
Calculated fit
Source
2

Phi-4 Mini

Microsoft

4B total · 4B active

INT4INT4

Foundry Local

75% less weight memory

4 GB unified memory

4 GB memory floor

Context not stated

loadable
No comparable result published
Calculated fit
Source
3

Qwen3 8B

Qwen

8B total · 8B active

INT4GGUF

llama.cpp

75% less weight memory

8 GB unified memory

8 GB memory floor

Context not stated

loadable
No comparable result published
Calculated fit
Showing 3 of 16
How hardware fit and quality loss are measured

Each row is a model revision, precision or quantization method, runtime, workload profile, and hardware configuration. Artifact or parameter-derived weight memory excludes KV cache and runtime overhead unless a configuration states otherwise. Quality deltas are published only when the same model baseline, benchmark, and evaluation setup are identifiable; missing comparisons remain explicitly unavailable.

Weight memory is separate from KV cache, activations, runtime overhead, context, and concurrency. A memory fit is therefore a floor, not a throughput promise.

Quality changes appear only for directly comparable baseline and quantized evaluations. Direction is normalized for lower-is-better metrics; when several comparisons exist, quality sorting uses their median. Missing results remain unpublished instead of being replaced by a technique-wide estimate.