willitrun·ai

Model comparison

Gemma 4 31B vs Gemma 4 12B

Gemma 4 31B vs Gemma 4 12B for local inference — VRAM, tokens/sec on RTX 4090 and M4 Max, and quality benchmarks side by side. Gemma 4 31B leads on quality, while Gemma 4 12B decodes faster on an RTX 4090.

MetricGoogle Gemma 4 31BGoogle Gemma 4 12B
Parameters30.7B12B
Active params (MoE)densedense
Context window256K262K
Quality tier8674
MMLU-Pro85.2
GPQA Diamond84.3
SWE-bench Verified
LiveCodeBench80.0
VRAM (Q4_K_M)18.7 GB7.3 GB
Speed — RTX 4090 (tok/s)13.0109.9
Speed — M4 Max (tok/s)25.543.4

Which should you run?

Gemma 4 31B scores higher on the quality benchmarks, so pick it when answer quality matters most. Gemma 4 12B is faster to decode on an RTX 4090 (110 vs 13 tok/s). Gemma 4 12B needs less VRAM (7.3 GB at Q4_K_M), so it fits on smaller GPUs.