Open-model infrastructure
LLM hardware requirements
Compare the hardware needed for downloadable models across precision, quantization method, runtime, workload profile, and source strength. Each rank represents one concrete configuration, not an absolute barrier.
- Data as of
- Aug 25, 2026
16
44
40
0
Hardware ladder
Model configurations by deployment footprint
Counts overlap when a model has several validated or calculated configurations.
Local
CPU / unified
6 models
1 GPU
Single accelerator
7 models
2–4 GPUs
Multi-GPU
5 models
GPU server
5–8 GPUs
9 models
Cluster
9+ GPUs
1 model
Showing published configurations.
Ranked configurations
16 models match
One primary configuration per model · alternatives expand in place
| Rank | Model | Variant | Hardware | Profile | Quality delta | Evidence |
|---|---|---|---|---|---|---|
| 1 | Gemma 4 E2B 5B total · 5B active | INT4INT4 vLLM 75% less weight memory | 4 GB unified memory 4 GB memory floor Context not stated | loadable | No comparable result published | Calculated fit |
| 2 | Phi-4 Mini Microsoft 4B total · 4B active | INT4INT4 Foundry Local 75% less weight memory | 4 GB unified memory 4 GB memory floor Context not stated | loadable | No comparable result published | Calculated fit |
| 3 | Qwen3 8B Qwen 8B total · 8B active | INT4GGUF llama.cpp 75% less weight memory | 8 GB unified memory 8 GB memory floor Context not stated | loadable | No comparable result published | Calculated fit |
How hardware fit and quality loss are measured
Each row is a model revision, precision or quantization method, runtime, workload profile, and hardware configuration. Artifact or parameter-derived weight memory excludes KV cache and runtime overhead unless a configuration states otherwise. Quality deltas are published only when the same model baseline, benchmark, and evaluation setup are identifiable; missing comparisons remain explicitly unavailable.
Weight memory is separate from KV cache, activations, runtime overhead, context, and concurrency. A memory fit is therefore a floor, not a throughput promise.
Quality changes appear only for directly comparable baseline and quantized evaluations. Direction is normalized for lower-is-better metrics; when several comparisons exist, quality sorting uses their median. Missing results remain unpublished instead of being replaced by a technique-wide estimate.