willitrun·ai

Prism ML

Ternary Bonsai 27B

Frontera
611.7KDescargas1.0KMe gustaJul 2026Publicado262K tokensContextoApache 2.0Licencia87 FuerteCalidad

Ternary Bonsai 27B (27B parameters) requires approximately 9.7 GB of VRAM with Q2_0_G128 quantization. For the best balance of quality and speed, we recommend hardware with at least 12 GB of VRAM.

Comenzar

— copia y pega para ejecutar en local

Copy-paste commands to run Ternary Bonsai 27B on your machine.

Run

docker run --rm -it ghcr.io/ggerganov/llama.cpp:full \ --hf-repo "prism-ml/Ternary-Bonsai-27B-gguf" \ --hf-file "Ternary-Bonsai-27B-gguf-Q2_0_G128.gguf" \ -c 4096 -ngl 99

Quick specs

Parameters27B
Architecturedense
Context262K tokens
Modalitytext+vision
Min RAM7.2 GB
Rec. RAM7.2 GB (Q2_0_G128)
LicenseApache 2.0
FamilyBonsai
Code Chat Reasoning

About this model

Ternary Bonsai 27B is Prism ML's ternary-weight build of Qwen3.6-27B, running full 27B-class reasoning at a true 1.71 bits per weight. It deploys in roughly 7.2 GB instead of ~54 GB at FP16 while retaining about 95% of full-precision quality, and keeps the 262K context practical via the Qwen3.6 hybrid-attention backbone and 4-bit KV-cache quantization.

  • ~7.2 GB deployed footprint for a 27B model — about 9.4x smaller than FP16.
  • Retains ~95% of FP16 quality: 80.49 average across 15 thinking-mode benchmarks.
  • True 1.71 bits/weight end-to-end — embeddings, attention, MLP and LM head, with no high-precision escape hatches.
  • Scores above a conventional IQ2_XXS build (72.73) at under two-thirds of its footprint.
  • ~26 tok/s on an Apple M5 Pro laptop; ships a DSpark drafter for a 1.34x CUDA speedup.

Modelos relacionados

Tu hardware

Detectando...

Selecciones rápidas

Mejor hardware

Mejores opciones para Ternary Bonsai 27B

Ejecutar este modelo

Opciones de cuantización

Estimaciones de VRAM por nivel de cuantización

No hardware detected — fit column shows raw VRAM estimates

QuantBitsVRAMQualityFit
Q1_0_G128
1.125
3.9 GB
Very Low
Q2_0_G128
1.71
7.2 GB
Low
Q2_K
2
10.5 GB
Low
Q3_K_S
3
13.2 GB
Low
NVFP4
4
15.1 GB
Medium
Q4_K_M
4
16.5 GB
Medium
Q5_K_M
5
19.4 GB
High
Q6_K
6
22.1 GB
High
Q8_0
8
28.9 GB
Very High
F16
16
55.4 GB
Maximum

Compatibilidad de hardware

Estimaciones de encaje en todo el hardware

Abrir calculadora

Computing compatibility...

Desglose de memoria

Reference: RTX 2060 6GB

Weights7.2 GB
KV Cache1.0 GB
Runtime0.9 GB
Headroom0.6 GB

Preguntas frecuentes

FAQ — Ternary Bonsai 27B

Ver también