Qwen
Qwen 3.5 9B
Warum empfohlen
This model is a direct match for coding. It belongs to a current frontier family for local AI. It fits natively with comfortable headroom. Known channels: huggingface, ollama, lm-studio.
Capacity: Roomy · Bandwidth: Medium · Stack: Standard
Interactive: Good · Light API: Great · Bottleneck: Balanced
Punktzahl
130.7
Passungsstatus
Runs well
Passung: Runs well mit sicherem Kontext 32K.
Laufzeit-Support: unknown via n/a auf unknown.
Laufzeit
llama.cpp
Artefakt
n/a
Quant.
Q4_K_M
Dekodierung
71.5 tok/s
Sicherer Kontext
32K
Offizieller Kontext
131K
Support
n/a
TTFT
2708 ms
Gewichte: 5.5 GB
KV-Cache: 2.2 GB
Backend: unknown
Current limits
This setup is broadly balanced for this model.
No major red flags
This recommendation has enough memory headroom and acceptable estimated speed for the selected workload.
Best next improvements
Punktzahl 130.7 kombiniert Workload-Übereinstimmung, Katalogaktualität, Passungssicherheit, Kontextabdeckung, Artefaktwahl, Speicherauslastung, Durchsatz und Latenz.