← Back to Models
Metamoe-transformer
Llama 4 Maverick
400B total / 17B active · load estimate uses total • 243.4GB VRAM (Q4 estimate) • 1,000,000 context
Estimate · Q4_K_M · 4k context · batch 1. For dense 7B–70B this sits ~5–8% above the Q4_K_M GGUF file. Not peak runtime VRAM and not a measured bench.
Specifications
Parameters
400B total / 17B active · load estimate uses total
VRAM (Q4 est.)
243.4 GB
VRAM (FP16)
800 GB
Context Window
1,000,000
Architecture
moe-transformer
License
Llama-4
Finding cloud alternatives...
Run Llama 4 Maverick on Mac Studio M3 Ultra 256GB
~243.4GB VRAM needed at Q4. Mac Studio M3 Ultra 256GB has 256GB — buy hardware or rent cloud.
We may earn a commission from cloud and hardware partners at no extra cost to you.
Find the cheapest GPU that can run Llama 4 Maverick
Compatibility Lab — check every GPU × quantization combination