AI startup PrismML has developed a groundbreaking compression technique that reduces a large language model to fit on everyday PCs and possibly smartphones, without significantly compromising its reasoning capabilities.
- Bonsai 2 27B compresses Qwen3.8 27B down to 5.9 GB for PC and mobile use
- Model retains 98% of original benchmark performance with ternary weight compression
- Next goal: compressing larger models with even better accuracy retention
What happened
PrismML, an AI startup founded by Caltech researchers, launched Bonsai 2 27B, a large language model that significantly compresses Alibaba’s Qwen3.8 27B from its original size to only 5.9 GB. This reduction makes it possible to run sophisticated LLMs on personal computers and potentially high-end smartphones. The company’s proprietary compression approach uses ternary weights, simplifying data storage within the model and achieving a 9-10x memory reduction.
This release marks a notable improvement over PrismML’s earlier model, which retained 95% benchmark accuracy, with Bonsai 2 now reaching 98%. The company has distributed millions of downloads, demonstrating strong user interest. CEO Babak Hassibi also hinted at forthcoming models containing hundreds of billions of parameters, which may compress even more efficiently due to scaling advantages.
Why it matters
The ability to compress and run advanced LLMs directly on personal devices addresses major hurdles in AI deployment, including reliance on cloud infrastructure, latency, cost, and data privacy concerns. PrismML’s models are designed to deliver near state-of-the-art natural language understanding and reasoning performance without needing high-end servers, potentially democratizing AI access.
Experts see this development as a potential game changer, especially as mobile and edge computing continue to grow. Running intelligence locally means users can benefit from faster responses, better privacy, and reduced dependency on internet connectivity. The company also benefits from backing by notable investors and advisors, lending credibility and resources to its efforts.
What to watch next
PrismML’s next key milestone will be successfully compressing significantly larger models beyond 27 billion parameters, targeting several hundred billion parameter scales. Due to the properties of larger models, the startup expects less intelligence loss during compression, which would make powerful AI even more versatile on consumer devices.
Additionally, negotiations and potential partnerships with major technology players, such as rumored talks with Apple, could accelerate adoption and integration of PrismML’s compressed models into mainstream consumer hardware. Monitoring how these relationships evolve and any new product announcements will provide insight into the future of on-device AI computing.