PrismML's new Bonsai 27B model, based on Qwen3.6 27B, introduces multi-step reasoning, structured tool use, and agentic loops to local AI deployments. It achieves this by offering two highly quantized variants: Ternary Bonsai 27B at 5.9 GB (95% performance retention) and 1-bit Bonsai 27B at 3.9 GB (90% performance retention). These variants make 27B-class models practical for local deployment, which previously required 18-54 GB, and are open-sourced under the Apache 2.0 license.
Extreme quantization enables 27B-class models to run on edge devices, shifting inference demand from cloud to local hardware and expanding the addressable market for on-device AI.