← Back to Models
OpenAImoe
gpt-oss 120B
117B total / 5.1B active · load estimate uses total • 71.2GB VRAM (Q4 estimate) • 128,000 context
Estimate · Q4_K_M · 4k context · batch 1. For dense 7B–70B this sits ~5–8% above the Q4_K_M GGUF file. Not peak runtime VRAM and not a measured bench.
Native weights ship as MXFP4. The Q4_K_M figure is the shared formula on total parameters — not the MXFP4 pack size.
Specifications
Parameters
117B total / 5.1B active · load estimate uses total
VRAM (Q4 est.)
71.2 GB
VRAM (FP16)
234 GB
Context Window
128,000
Architecture
moe
License
Apache-2.0
Finding cloud alternatives...
Run gpt-oss 120B on RTX PRO 5000 72 GB Blackwell
~71.2GB VRAM needed at Q4. RTX PRO 5000 72 GB Blackwell has 72GB — buy hardware or rent cloud.
We may earn a commission from cloud and hardware partners at no extra cost to you.
Find the cheapest GPU that can run gpt-oss 120B
Compatibility Lab — check every GPU × quantization combination