On-Device AI: What It Means and Why It Is Suddenly Possible

Almost every AI product you use sends your input to a data center. You type, it goes over the network, a very large model on very expensive hardware produces a response, and the response comes back. That architecture is why these products need subscriptions, why they stop working on a plane, and why your conversations exist on somebody’s servers.

On-device AI means the model runs on the hardware in your hand. No network call, no server, no account. The same category of technology, in a completely different shape.

Five years ago this was not realistic for language models. Now it is, and the reasons why are interesting.

What actually changed #

Three things happened at roughly the same time, and none of them alone would have been enough.

Small models got dramatically better. A model with a billion parameters today outperforms a model with ten times that many from a few years ago. The improvement came from better training data, better training methods, and from distillation, where a large model is used to train a small one to imitate its behavior. The capability floor moved up faster than the size requirement moved down.

Quantization got good. Models store their weights as numbers, and those numbers were traditionally 16 or 32 bits each. Quantization stores them in 4 or 8 bits instead, cutting memory use by a factor of four or eight with surprisingly little quality loss. This is the single biggest reason a model fits on a phone at all, and it has its own explainer.

And phone hardware caught up. Modern phone chips include dedicated neural processing units, and memory bandwidth improved substantially. A current flagship phone has more machine learning throughput than a mid-range desktop GPU from a few years ago.

Multiply those together and a model that would have needed a server in 2020 now runs on a device in your pocket.

The four constraints #

Building for on-device means designing around four hard limits. Every decision traces back to one of them.

Memory. A phone has maybe 6 to 12 GB of RAM, shared with the operating system and everything else running. The model has to be fully loaded into memory to run. Exceed the budget and the OS kills your app without warning or ceremony. This is the binding constraint, and it is why on-device models top out where they do.

Storage and download size. Nobody downloads 20 GB to try an app. Realistically you have a few hundred megabytes to a few gigabytes, and the download is the first thing a user experiences, which makes it a product problem as much as a technical one.

Speed. Generation speed is measured in tokens per second, and below a certain rate it stops feeling like a conversation and starts feeling like waiting. The threshold where people lose patience is around ten to fifteen tokens per second for chat.

Heat and battery. Sustained inference runs the chip hard. Phones thermally throttle, which means a model that is fast for thirty seconds gets slower after two minutes. And a feature that visibly drains the battery is a feature people turn off.

Server-side AI has none of these problems, which is precisely why almost everyone builds server-side.

What a small model can and cannot do #

Being honest about this matters, because the fastest way to lose someone is to overpromise.

What works well on-device: summarizing text you provide, rewriting and tone adjustment and editing, answering general knowledge questions where being approximately right is fine, drafting emails and messages and short documents, brainstorming, translation between common languages, classification and extraction from text, and transcription, which is a genuinely different case covered below.

What does not work well on-device: long, multi-step reasoning, since smaller models lose the thread. Anything requiring deep factual precision, since small models hallucinate more than large ones and confidently. Complex code generation across multiple files. Very long documents, since on-device context windows are usually smaller. And anything needing current information, because the model is frozen and has no network.

A good on-device model is roughly a capable assistant for everyday tasks, and it is not a frontier model. It will not write your dissertation. It will draft your email on a plane with no signal, which the frontier model cannot do at all.

The case where on-device clearly wins #

There are three, and they are worth separating because they are different arguments.

Privacy that is architectural rather than promised. Every cloud AI service has a privacy policy. Policies are commitments a company makes and can change, and they depend on the company’s practices, its subprocessors, and its future ownership. A model running locally with no network permission makes a different kind of statement: the data cannot leave, because there is no path for it to leave. That is verifiable in a way a policy is not.

This matters most for categories where people genuinely will not use a cloud tool: medical conversations, legal matters, journalism and sources, therapy notes, confidential business material. The market for these is not everyone. It is people for whom the cloud version is a non-starter.

Working offline. Planes, subways, rural areas, foreign countries without a data plan, and any situation where connectivity is unreliable. A cloud AI app is a blank screen without a network.

Cost. Running inference on a server costs money per request, forever, which is why cloud AI is a subscription. On-device inference costs the developer nothing, which means the app can be free without a business model that involves you.

Speech is the sleeper case #

Language models get the attention, but speech-to-text is arguably a better fit for on-device.

Speech recognition models are much smaller than language models. Whisper, the open model most on-device transcription is built on, comes in sizes from about 75 MB up to a few gigabytes. The smaller ones run comfortably on any modern phone.

And the privacy argument is sharper for audio than for text. A recording of a doctor’s appointment, a therapy session, an interview with a source, or a business call is more sensitive than most typed queries, and cloud transcription services upload all of it, frequently retain it, and in some cases use it to train future models.

I built Private Transcribe around exactly this: Whisper running locally, 99 or more languages, and audio that never reaches a network. The build write-up covers the engineering side, and how speech recognition actually works covers the model side.

For chat, I built Personal LLM, which runs small quantized models on the phone with no account and no subscription. Same write-up treatment if you want the details of what was hard about it.

When on-device is the wrong answer #

To be useful this has to cut both ways.

If you need frontier capability, use the cloud. A 3-billion-parameter model on a phone is not going to match a frontier model, and pretending otherwise wastes your time. If you need current information, use something with web access. If you need to process a two-hundred-page document, the context window on a phone will not hold it.

On-device does not replace cloud AI. There are tasks where local is clearly better, tasks where cloud is clearly better, and a growing middle where local is good enough and free and private, which makes it the default choice.

That middle is what has been expanding, and it will keep expanding, because small models improve faster than large ones in relative terms and phone hardware keeps getting better.