PrismML Shrinks 56GB AI Model to 5.9GB
PrismML has launched Bonsai 2 27B, a compact multimodal AI model that fits on consumer PCs and high-end mobile devices. Using ternary compression, the company reduced Qwen3.8's 56GB footprint to just 5.9GB while retaining 98.2% of its capabilities, outperforming standard quantization methods that typically sacrifice accuracy.
The model runs on Nvidia GPUs and Apple devices including iPhone and iPad, achieving 143 tokens per second on an RTX 5090. It is 40% more energy-efficient than comparable 8B models and enables fully local AI inference, eliminating cloud dependency and protecting user privacy. Weights are available under an Apache 2.0 license.
