PrismML drops ternary Bonsai 2 27B
PrismML has released Ternary Bonsai 2 27B, a compressed open-weights multimodal model based on Qwen3.8 27B that delivers 27B-class intelligence within a 5.9GB memory footprint—over 9x smaller than its full-precision counterpart. Using ternary weights with FP16 group-wise scaling to reach 1.76 effective bits per weight, the Apache 2.0-licensed model retains 98.2% of full-precision benchmark performance across reasoning, coding, vision, and agentic tasks.
Near-lossless sub-2-bit ternary compression makes 27B-class multimodal intelligence a viable, fast reality for consumer edge devices and private local agentic workflows.
• Quantization below 4 bits historically suffered steep capability cliffs in multi-step agentic workflows and coding; closing the benchmark gap to 98.2% effectively renders compression lossless for practical applications.
• Operating in only 5.9GB of memory unlocks local execution of high-tier multimodal reasoning on everyday consumer hardware, laptops, and mobile devices without requiring expensive workstation VRAM.
• Sustained throughputs of up to 143 tokens/second combined with high energy efficiency make real-time coding loops and computer-use agents battery-friendly and latency-competitive with cloud APIs.
• Elevating "intelligence density per gigabyte" highlights a shifting paradigm where optimized low-bit weights and custom hardware kernels challenge the economic assumptions of both edge and datacenter inference.
DISCOVERED
1h ago
2026-09-18
PUBLISHED
4h ago
2026-09-17
RELEVANCE
AUTHOR
JonSchneider