willitrun·ai

Prism ML

1-bit Bonsai 27B

Frontier
2.1MDownloads633CurtidasJul 2026Publicado262K tokensContextoApache 2.0Licença82 ForteQualidade

1-bit Bonsai 27B (27B parameters) requires approximately 6.4 GB of VRAM with Q1_0_G128 quantization. For the best balance of quality and speed, we recommend hardware with at least 8 GB of VRAM.

Comece agora

— copie e cole para rodar localmente

Copy-paste commands to run 1-bit Bonsai 27B on your machine.

Run

docker run --rm -it ghcr.io/ggerganov/llama.cpp:full \ --hf-repo "prism-ml/Bonsai-27B-gguf" \ --hf-file "Bonsai-27B-gguf-Q1_0_G128.gguf" \ -c 4096 -ngl 99

Quick specs

Parameters27B
Architecturedense
Context262K tokens
Modalitytext+vision
Min RAM3.9 GB
Rec. RAM3.9 GB (Q1_0_G128)
LicenseApache 2.0
FamilyBonsai
Code Chat Reasoning

About this model

1-bit Bonsai 27B is Prism ML's binary-weight build of Qwen3.6-27B, packing a full 27B-class reasoning model into a true 1.125 bits per weight. It deploys in roughly 3.9 GB instead of ~54 GB at FP16 — putting a 27B model on everyday laptops and entry-level GPUs — while retaining about 90% of full-precision quality.

  • ~3.9 GB deployed footprint for a 27B model — about 14.2x smaller than FP16.
  • Retains ~90% of FP16 quality: 76.11 average across 15 thinking-mode benchmarks.
  • True 1.125 bits/weight end-to-end; the vision tower ships in compact 4-bit HQQ.
  • Keeps thinking, reasoning and agentic behaviour in the sub-4-bit regime where conventional low-bit builds collapse.
  • ~44 tok/s on an Apple M5 Pro laptop, via custom 1-bit llama.cpp kernels (CUDA, Metal).

Modelos relacionados

Seu hardware

Detectando...

Escolhas rápidas

Melhor hardware

Melhores opções para 1-bit Bonsai 27B

Rodar este modelo

Opções de quantização

Estimativas de VRAM por nível de quantização

No hardware detected — fit column shows raw VRAM estimates

QuantBitsVRAMQualityFit
Q1_0_G128
1.125
3.9 GB
Very Low
Q2_0_G128
1.71
7.2 GB
Low
Q2_K
2
10.5 GB
Low
Q3_K_S
3
13.2 GB
Low
NVFP4
4
15.1 GB
Medium
Q4_K_M
4
16.5 GB
Medium
Q5_K_M
5
19.4 GB
High
Q6_K
6
22.1 GB
High
Q8_0
8
28.9 GB
Very High
F16
16
55.4 GB
Maximum

Compatibilidade de hardware

Estimativas de compatibilidade para todo o hardware

Abrir calculadora

Computing compatibility...

Detalhamento de memória

Reference: RTX 2060 6GB

Weights3.9 GB
KV Cache1.0 GB
Runtime0.9 GB
Headroom0.6 GB

Perguntas frequentes

FAQ — 1-bit Bonsai 27B

Veja também