TechNewsReel
Live

PrismML Shrinks 27B Model to 3.9GB for Smartphone Deployment

Bonsai 27B uses extreme quantization to fit a high-parameter reasoning model on edge devices like the iPhone 17 Pro.

TechNewsReel Newsroom · August 17, 2026

PrismML has released Bonsai 27B, a highly compressed version of the Qwen3.6-27B model designed to run locally on edge devices. The release marks a significant technical milestone in fitting high-parameter reasoning models onto consumer hardware, including smartphones.

Rather than being a new pre-trained model, Bonsai 27B is a low-bit representation of the Qwen3.6-27B architecture. By employing 1-bit and ternary quantization, PrismML has drastically reduced the model's memory footprint. The 1-bit variant, which utilizes binary weights, occupies just 3.9GB, while the ternary variant—using {-1, 0, +1} weights—occupies 5.9GB. For comparison, the original 16-bit version of the model requires 54GB of memory. According to PrismML, this makes Bonsai 27B the first model of its capability class capable of running on a phone, specifically the iPhone 17 Pro.

The Push for Intelligence Density

This development is part of a broader effort by PrismML to maximize "intelligence density," or the amount of capability available per gigabyte of memory. Founded by researchers from Caltech and backed by Google and Khosla Ventures, the company is focusing on enabling local, private, and cost-effective agentic AI.

Traditionally, 27B parameter models have been inaccessible to mobile users; even standard 4-bit quantization typically requires roughly 18GB of VRAM, which exceeds the capacity of most laptops and nearly all smartphones. By pushing quantization to the 1-bit and ternary levels, PrismML removes the hardware barrier for models of this scale.

Trade-offs in Performance

While the memory reduction is drastic, the high parameter count creates a bottleneck in generation speed. Serdar Yegulalp of InfoWorld noted that despite its compactness, Bonsai is not the fastest model, observing that its 27 billion parameters resulted in token output that was generally slower than smaller models, such as the Qwen 7B.

Real-world testing on consumer-grade hardware highlights this tension between size and speed. On an RTX 5060 with 8GB of VRAM, the model averaged between 10 and 20 tokens per second, though it reached a peak of 40 tokens per second. This suggests that while the model fits in memory, the computational overhead of processing 27 billion parameters remains a limiting factor for real-time interaction.

The Future of On-Device AI

The ability to run a 27B-class model on a mobile device signals a shift toward persistent, on-device agentic AI. By eliminating the need for cloud connectivity, developers can build assistants capable of multi-step reasoning and tool-calling without the associated API costs, cloud latency, or privacy risks of transmitting sensitive data to remote servers.

Industry observers will now be watching to see if further optimization can resolve the token generation lag and whether other large-scale models will adopt similar extreme quantization techniques to compete for the edge computing market.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.