Fit estimates and live market context for Llama 8B through 70B. Same Q4_K_M · 4k · batch 1 helper as the VRAM calculator.
Llama 3.1 70B Q4 is ~42.6 GB. It does not fit a 24 GB card. Numbers below are estimates, not benches.
24 GB catalog VRAM. Consumer reference for 8B-class Q4. 70B Q4 does not fit this card.
Affiliate links: we may earn a commission at no extra cost to you.
Catalog VRAM only. No invented throughput or street prices.
| Model | FP16 estimate | Q8 estimate | Q4_K_M estimate |
|---|---|---|---|
| Llama 3.1 8B | 17.3 GB | 8.7 GB | 4.9 GB |
| Llama 3.1 70B | 151.3 GB | 76.3 GB | 42.6 GB |
| Llama 3.1 405B | 875.4 GB | 441.7 GB | 246.5 GB |
Estimate · Q4_K_M · 4k context · batch 1. For dense 7B–70B this sits ~5–8% above the Q4_K_M GGUF file. Not peak runtime VRAM and not a measured bench.
Same Q4_K_M · 4k · batch 1 helper as the VRAM calculator. A 24 GB card does not hold a 70B-class Q4 row (~42.6 GB).
Open VRAM calculatorOpen the live cloud GPU tracker. Prices come from the current feed — not a fixed dollar floor. Catalog provider count and live offering count are labeled separately.
Open live cloud tracker