Run Google Gemma models locally. GPU picks for 2B through 27B variants.
16GB VRAM • Runs Gemma 2 27B at Q4 • 40+ tok/s on 9B
Affiliate links: we may earn a commission at no extra cost to you.
| Model | Full Precision | Q8 (8-bit) | Q4 (4-bit) |
|---|---|---|---|
| Gemma 2 2B | 4 GB | 2.5 GB | 1.5 GB |
| Gemma 2 9B | 18 GB | 10 GB | 6 GB |
| Gemma 2 27B | 54 GB | 29 GB | 16 GB |
* Add 1-2GB overhead for context window. Values are approximate.
Check if your GPU can run specific Gemma models at every quantization level.
Open VRAM CalculatorRent GPU compute from $0.39/hr. Compare 24+ providers with live pricing.
Browse Cloud GPUs