Friday, September 18, 2026
Tech Beat
Sep 17, 2026, 10:34 PMArtificial Intelligence

PrismML Shrinks Alibaba’s Qwen3.8 27B to 5.9 GB for Local AI

PrismML shrinks Alibaba’s Qwen3.8 27B to 5.9 GB while retaining 98% of benchmark performance, targeting private AI on PCs and some high-end smartphones.

Listen to this briefingAudio briefing

Summary

On Thursday, September 17, 2026, PrismML released Bonsai 2 27B, compressing Alibaba’s open source Qwen3.8 27B to 5.9 GB, a ninefold to tenfold memory reduction that could fit PCs and premium smartphones. Bonsai 2 retains 98% of Qwen’s aggregate benchmark scores, versus 95% for the first Bonsai released in March. PrismML says that model has surpassed 11 million downloads, while smaller variants have added 2.6 million.

PrismML reduces normally 16 bit model weights to ternary values of plus one, minus one or zero. Perfect parity remains unproven, although the 2% benchmark gap may have limited practical effect because base models and benchmarks are imperfect and surrounding software also shapes accuracy. CEO Babak Hassibi, a Caltech professor and compression expert, expects models with several hundred billion parameters within the next couple of months, arguing larger models may retain intelligence more easily when compressed.

Founded by Caltech researchers, PrismML has raised a $22.25 million seed round and is backed by Khosla Ventures, Cerberus Capital and Caltech. Adviser Ion Stoica co-founded Databricks and directs Berkeley’s Sky Computing Lab. PrismML competes with Multiverse Computing, founded by a Donostia International Physics Center professor. Hassibi declined to confirm rumored Apple talks. Local deployment could make advanced AI free to run on existing devices and keep user data out of the cloud.

Positives

  • Bonsai 2 reduces Qwen3.8 27B to 5.9 GB while preserving 98% of its aggregate benchmark performance.
  • Memory requirements fall ninefold to tenfold, potentially bringing reasoning models to PCs and premium smartphones.
  • The first Bonsai has exceeded 11 million downloads, while PrismML’s smaller models have added 2.6 million.
  • Ternary weights reduce normal 16 bit values to plus one, minus one or zero.
  • Local models could run without cloud fees while keeping private data on users’ devices.

Risks & concerns

  • Bonsai 2 still trails Qwen3.8 27B by 2% on aggregate benchmarks, and Hassibi expects compression to retain some performance cost.
  • Perfect benchmark parity remains unproven despite improvement from 95% with the first Bonsai to 98% with Bonsai 2.
  • Compression of models with several hundred billion parameters has not yet been demonstrated, with releases only targeted within the next couple of months.
  • Multiverse Computing is also pursuing LLM compression, creating competition for PrismML’s $22.25 million funded effort.
  • Rumored discussions with Apple remain unconfirmed because Hassibi declined to comment.
Primary sourceTechCrunchhttps://techcrunch.com/2026/09/17/prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceSep 17

Google DeepMind Institute Proposes AGI Tests and Possible Slowdown

Artificial IntelligenceSep 17

OpenAI's 12,000-Agent Hugging Face Incident Fuels AI Monitoring Boom

Artificial IntelligenceSep 17

OpenAI Finds GPT-5.6 Sol Hiding Misalignment From Successor Agents