How to Run Ternary-Bonsai-27B Locally — Setup Guide (August 2026)

Ternary-Bonsai-27B is a 27B model quantized down to a 1.7-bit ternary format — small enough on disk that a 16 GB card can hold it with a huge context window instead of a squeezed one. It needs a custom llama.cpp fork, since its ternary kernels aren't in stock llama.cpp, which makes it a good showcase of TurboLLM's actual differentiator: run the fork a model actually needs, without compiling it yourself.

Numbers you can trust

We don't print guessed speeds. Every measured number on this page came from TurboLLM's own auto-tuner running Ternary-Bonsai-27B on real hardware — labeled as such. When you run it yourself, TurboLLM benchmarks it on your exact GPU and shows a VRAM-fit verdict before you load, and the real measured tokens/sec once it's running.

Does it fit your GPU?

It's a reasoning model (emits a visible thinking step) with vision input support, quantized with a 1-bit-class ternary scheme rather than the usual 4/5/8-bit GGUF quants.

Ternary-Bonsai-27B (Q2_0, ternary)
community (PrismML) · GGUF
Dense · 1.7-bit ternary · vision + reasoning
thinkingvision
VRAMFit
16 GBQ2_0 (~7.2 GB on disk) + q8_0 KV cache → fits 15.7 / 16 GB at a full 200K context
prism-ml/Ternary-Bonsai-27B-gguf

How to run it

  1. Install TurboLLM

    npx turbollm — no install step, works on Windows, macOS, and Linux. It detects your GPU and auto-provisions a matching llama-server build (no CUDA toolkit, no Python, no compiler).

  2. Get the model

    Paste prism-ml/Ternary-Bonsai-27B-gguf into TurboLLM's in-app Hugging Face search — it lists every quant with a VRAM-fit verdict against your real free VRAM, and downloads are resumable and SHA-256 verified.

  3. Let auto-fit pick the setup

    Flip the auto-fit toggle and TurboLLM chooses the GPU/CPU layer split (and, for a Mixture-of-Experts model, the expert-offload split) for your card — the same manual flag-hunting every other guide walks you through by hand.

  4. Load it

    TurboLLM benchmarks the model on your exact GPU as it loads and shows the real measured tokens/sec — not a number copied from someone else's card.

  5. Use it

    Chat in the built-in UI, or point any OpenAI- or Anthropic-compatible tool (including Claude Code) at TurboLLM's local API — see the API docs.

Need the PrismML llama.cpp (branch prism) fork instead?

Ternary-Bonsai-27B's Q2_0 ternary kernels only exist in PrismML-Eng's fork of llama.cpp. Paste the fork's git URL into TurboLLM's "Add via git repo" flow and it builds automatically — about 3.5 minutes on Windows with CUDA, no manual cmake or toolchain wrangling. The KV-cache type matters as much as the quant here: switching from f16 to q8_0 KV is the difference between the model spilling to system RAM (~6.6 t/s) and staying fully resident (~62 t/s).

Measured, not guessed

RTX 5070 Ti 16 GBTurboLLM (measured)
Q2_0 + q8_0 KV cache, flash attention on, 200K ctx~62 t/s
Same model, f16 KV cache (spills to system RAM)~6.6 t/s

Measured in TurboLLM (app-displayed generation speed) on an RTX 5070 Ti 16 GB, July 2026.

FAQ

Why does Ternary-Bonsai-27B suddenly get slow or stutter?

The usual cause is silent VRAM spill: once a model's weights plus its KV cache stop fitting in VRAM, the driver quietly offloads the overflow to system RAM over PCIe, and that path runs 5–10× slower with no error message — just a sudden slowdown. TurboLLM's auto-fit sizes the quant and GPU/CPU split against your card's real free VRAM (including KV-cache growth at your chosen context) specifically to avoid this, and shows a VRAM-fit verdict before you load rather than after it's already crawling. Full mechanics in VRAM spill explained.

Which quant should I use for Ternary-Bonsai-27B?

Pick the highest quant that fits your VRAM with room left over for the KV cache and compute buffer — not just the model weights. As a starting point, Q4_K_M is usually the sweet spot (small quality loss for a large size drop); step up to Q5_K_M/Q6_K/Q8_0 if you have headroom, or down to Q3_K_M/an IQ quant if you don't. TurboLLM shows a VRAM-fit verdict for every quant against your card's real free VRAM before you download — see Quantization explained for what the letters and numbers actually mean, or llama.cpp flags that actually matter for the other settings worth knowing.

Will Ternary-Bonsai-27B run on my GPU?

Check the VRAM table above for the tier closest to your card. If you're not sure which tier you're in, or want picks tailored to your exact hardware, use What can I run? — enter your VRAM (or Apple unified memory) and it suggests the model, quant, and offload split that fits, the same logic TurboLLM's auto-fit runs when you load a model.

Run it now

$ npx turbollm

One command detects your GPU, provisions the right engine, and opens the UI. New here? Start with Install & first run and Quantization explained. On different hardware, see RTX 5070 Ti. Also see how to run it.