Run Microsoft Phi models locally. Efficient small models that shine on consumer GPUs.
12GB VRAM • Runs Phi-4 14B at Q4 • 20-30 tok/s • Best budget pick
Affiliate links: we may earn a commission at no extra cost to you.
| Model | Full Precision | Q8 (8-bit) | Q4 (4-bit) |
|---|---|---|---|
| Phi-3 Mini 3.8B | 8 GB | 4.5 GB | 2.5 GB |
| Phi-4 14B | 28 GB | 15 GB | 9 GB |
* Add 1-2GB overhead for context window. Values are approximate.
Check if your GPU can run Phi-4 at every quantization level.
Open VRAM CalculatorRent GPU compute from $0.39/hr. Compare 24+ providers with live pricing.
Browse Cloud GPUs