The India Decade

Caltech spin-off PrismML shrinks massive AI models to run on your phone

The startup’s new Bonsai 2 model compresses Alibaba’s 27-billion-parameter LLM by nearly 90 per cent while keeping its performance almost fully intact.

By The India Decade

Published

close up of a modern smartphone screen displaying complex code and mathematical symbols (file image)
close up of a modern smartphone screen displaying complex code and mathematical symbols (file image) · “Smartphone display screen” by Skitterphoto (CC0) via Wikimedia Commons

What happened

The dream of running powerful, human-like artificial intelligence directly on your laptop or phone—without sending your personal data to a remote server—has long been held back by a simple physical reality: these models are just too big.

A new release from Caltech spin-off PrismML suggests that barrier is falling. On Thursday, the startup launched Bonsai 2 27B, an ultra-compressed version of Alibaba's popular Qwen3.8 27B open-source model. By stripping away redundant data, PrismML has shrunk the model’s footprint down to just 5.9 gigabytes—a nine- to ten-fold reduction in memory.

The result is a model small enough to run comfortably on a standard PC and potentially on high-end smartphones, without requiring an active internet connection.

How it works

PrismML achieves this drastic shrinkage by changing how the model stores what it has learned. Standard large language models represent their internal "weights"—the parameters that dictate how they respond to prompts—using 16 bits of data per weight. PrismML uses a technique called "ternary" weights, which simplifies those values into just three possibilities: +1, -1, or 0. This radically reduces the amount of storage space each weight requires.

While shrinking a model usually degrades its accuracy, PrismML’s shrunken model retains 98 per cent of the original's aggregate benchmark scores. That is a notable improvement over PrismML's first release in March, which retained 95 per cent of the original's performance. That early version has already racked up more than 11 million downloads, alongside 2.6 million downloads for the company's even smaller models.

PrismML’s CEO Babak Hassibi, a Caltech professor and compression specialist, believes the trade-off is becoming negligible. In real-world use, a two per cent drop in benchmark performance is almost imperceptible to users.

Why this matters

Running advanced models directly on user hardware changes the economics and privacy of AI. "You are going to have intelligence at your fingertips," said Ion Stoica, a prominent computer scientist, Databricks co-founder, and adviser to PrismML. "It’s going to be free because it’s going to run on the device you already bought. It’s also going to be private, because you’re not going to send it to the cloud."

PrismML is operating in a competitive space, with rivals like the heavily funded Multiverse Computing also working on model compression. But PrismML’s academic pedigree has already turned heads. Alongside Stoica’s advisory role, the company has raised $22.25 million in a seed round backed by Khosla Ventures, Cerberus Capital, and Caltech.

There are also persistent industry rumours that the startup is in discussions with Apple, though Hassibi declined to comment on those talks.

What happens next

The startup's next move is to apply the same compression technique to much larger models in the hundreds-of-billions of parameters range. Paradoxically, Hassibi says this task should be easier. With larger models, there is simply more mathematical "room" to compress the data without damaging the underlying intelligence. He expects to release these larger compressed models within the next couple of months.

Found an error in this story? Write to editor@theindiadecade.com. Our corrections policy explains how we put mistakes right.

Key numbers

Bonsai 2 27B model size
5.9 GB
Source: PrismML
Seed funding raised
$22.25 million
Source: PrismML
Benchmark performance retention
98%
Source: PrismML
Original Bonsai model downloads
Over 11 million
Source: PrismML

In this story

  • Babak Hassibi — CEO of PrismML and professor at Caltech.
  • PrismML — The AI startup developing the model compression technology.

Topics

The day’s reporting, once each morning. No advertising.

More on this story