Text / Prism ML
Bonsai 2 · 27B
A compressed 27B-class reasoning model with publisher-provided GGUF packs and a dedicated llama.cpp fork for CUDA, Metal, and CPU.
Verified sourceReleased Sep 17, 2026Source checked 9/23/2026
LOCALRENTED GPUOPEN WEIGHTS
At a glance
- Parameters
- 27.36B
- Architecture
- Ternary hybrid attention
- Context length
- 262,144
- License
- Apache 2.0
- Disk space
- 7.21 GB
- Software
- PrismML llama.cpp fork
Sources and files
Publisher model card and package measurements Publisher runtime and pinned releasesWill it run on your machine?
Save your machine to see a personalized rating and its reasoning.
Add my hardwareBest for
- Private reasoning and coding on hardware that cannot hold conventional 27B FP16 weights.
Tradeoffs
- Requires Prism ML’s modified llama.cpp runtime; stock builds do not support these packed weights.
- The optional vision tower increases memory use for image input.
Ways to run it
PrismML llama.cpp fork · Windows, macOS, Linux
Use the publisher fork or its pinned binary. Stock llama.cpp cannot run the ternary packs correctly.
hf download prism-ml/Ternary-Bonsai-2-27B-gguf Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir .Explore its uses
REASONINGCODINGCOMPRESSED WEIGHTS
Get it running
Run a compressed private reasoning assistant
Use Bonsai 2’s dedicated low-bit runtime, download its PQ2_0 package, and verify a local response without accidentally using stock llama.cpp.
View workflowKeep exploring
Text
DeepSeekDeepSeek V4.1 Flash
A multimodal reasoning model with a compressed key-value cache, published as open weights.
Text
QwenQwen3.8-Flash-Next
An experimental open-weight multimodal model with sparse attention and 262K native context.
MODEL SIZE125B language · 6B active
Text
QwenQwen3 8B
An open-weight language model with switchable thinking and non-thinking modes.
MODEL SIZE8.2B