willitrun·ai

Model comparison

Gemma 4 31B vs Qwen 3.6 35B A3B

Gemma 4 31B vs Qwen 3.6 35B A3B for local inference — VRAM, tokens/sec on RTX 4090 and M4 Max, and quality benchmarks side by side. Qwen 3.6 35B A3B leads on quality, while Qwen 3.6 35B A3B decodes faster on an RTX 4090.

MetricGoogle Gemma 4 31BAlibaba Qwen 3.6 35B A3B
Parameters30.7B35B
Active params (MoE)dense3B
Context window256K262K
Quality tier8698
MMLU-Pro85.285.2
GPQA Diamond84.386.0
SWE-bench Verified73.4
LiveCodeBench80.080.4
VRAM (Q4_K_M)18.7 GB21.3 GB
Speed — RTX 4090 (tok/s)13.034.1
Speed — M4 Max (tok/s)25.543.7

Which should you run?

Qwen 3.6 35B A3B scores higher on the quality benchmarks, so pick it when answer quality matters most. Qwen 3.6 35B A3B is faster to decode on an RTX 4090 (34 vs 13 tok/s). Gemma 4 31B needs less VRAM (18.7 GB at Q4_K_M), so it fits on smaller GPUs.