Tech Behind ThingsHow the ordinary machinery actually works

Software

Why some machine learning runs on your device and some cannot

The split is decided by model size, memory bandwidth and power, and each of those constraints has a hard edge.

Close-up of colorful programming code on a blurred computer monitor.
Photograph by Al Nahian via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

These are listed in the order worth acting on, which with on-device inference is not the order they are usually presented in.

What matters most

  • Model weights must fit in memory before anything can run.
  • Generating output is limited by memory bandwidth more than by arithmetic.
  • On-device processing avoids sending data anywhere, which is a real privacy difference.

Weights have to fit

A model is a large table of numbers, and running it requires those numbers to be in memory the processor can reach. Reducing the precision of each number, a process called quantisation, shrinks the requirement several-fold with a modest loss of quality.

That is why the same model appears in versions that fit a phone and versions that need a server, with the smaller ones measurably less capable. When a model does not fit, no amount of patience makes it run; it simply cannot start.

Memory bandwidth is the usual bottleneck

Generating each unit of output requires reading a large share of the weights from memory. The processor therefore spends most of its time waiting for data rather than computing, which is why raw arithmetic capability is a poor predictor of speed.

Devices with faster unified memory run these workloads far better than their processor specifications suggest. This is also why adding more compute without more bandwidth produces disappointing gains.

Dedicated accelerators change the power equation

Neural processing units execute the specific operations these models need at far better performance per watt than a general processor. On a battery-powered device the constraint is energy per operation rather than peak speed, so that efficiency is what makes a feature viable.

At the protocol level, accelerators are specialised, so a model using operations they do not implement falls back to slower general execution. That fallback is invisible to users and explains why two similar features perform very differently on the same hardware.

What is kept on device, and why

Wake word detection, keyboard prediction, face grouping in photo libraries, live transcription and image processing are typically local. They are small models, they must run with low latency, and they handle data people would object to uploading. Large general-purpose models are typically remote because they exceed device memory by a wide margin.

In practice, hybrid designs run a small local model and escalate to a server for harder requests, which is why some responses are instant and some are not.

The privacy difference is genuine but not automatic

Processing locally means the raw data need not leave, which is a meaningful reduction in exposure. It does not prevent the result, or telemetry about the interaction, from being sent. Federated approaches train on many devices and share only model updates, which reduces but does not eliminate what can be inferred.

Mechanically, the honest question is what leaves the device, not where the computation happened.

Reading the claims

Parameter counts describe size, not capability, and comparisons across model families on that number alone are close to meaningless. Context length describes how much input a model can consider at once and has a direct cost in memory. Benchmark scores are sensitive to how questions are asked and to whether test material leaked into training data.

In practice, for a device feature, the useful questions are whether it works offline and how it behaves when it cannot reach a server.

Everything above, in order of what to do first

  1. Weights have to fit. A model is a large table of numbers, and running it requires those numbers to be in memory the processor can reach.
  2. Memory bandwidth is the usual bottleneck. Generating each unit of output requires reading a large share of the weights from memory.
  3. Dedicated accelerators change the power equation. Neural processing units execute the specific operations these models need at far better performance per watt than a general processor.
  4. What is kept on device, and why. Wake word detection, keyboard prediction, face grouping in photo libraries, live transcription and image processing are typically local.
  5. The privacy difference is genuine but not automatic. Processing locally means the raw data need not leave, which is a meaningful reduction in exposure.
  6. Reading the claims. Parameter counts describe size, not capability, and comparisons across model families on that number alone are close to meaningless.

The takeaway

It runs locally when it fits in memory and the power budget. Everything else is somebody else's computer.

The constraint is almost always physical, and marketing rarely mentions which one.

Questions readers ask

Does on-device processing mean nothing is uploaded?

It means the raw input need not be. Results, usage telemetry and error reports may still be sent. Check what the settings actually control rather than assuming.

Why is a feature fast on one phone and slow on another?

Usually memory bandwidth and whether the model can run on the dedicated accelerator. A fallback to general-purpose execution is often several times slower.

Softwaremachine learninginferenceacceleratorsprivacy
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida