willitrun·ai

Model comparison

Gemma 4 12B vs Gemma 4 31B

Gemma 4 12B vs Gemma 4 31B for local inference — VRAM, tokens/sec on RTX 4090 and M4 Max, and quality benchmarks side by side. Gemma 4 31B leads on quality, while Gemma 4 12B decodes faster on an RTX 4090.

MetricGoogle Gemma 4 12BGoogle Gemma 4 31B
Parameters12B30.7B
Active params (MoE)densedense
Context window262K256K
Quality tier7486
MMLU-Pro85.2
GPQA Diamond84.3
SWE-bench Verified
LiveCodeBench80.0
VRAM (Q4_K_M)7.3 GB18.7 GB
Speed — RTX 4090 (tok/s)109.913.0
Speed — M4 Max (tok/s)43.425.5

Which should you run?

Gemma 4 31B scores higher on the quality benchmarks, so pick it when answer quality matters most. Gemma 4 12B is faster to decode on an RTX 4090 (110 vs 13 tok/s). Gemma 4 12B needs less VRAM (7.3 GB at Q4_K_M), so it fits on smaller GPUs.