Tech Behind ThingsHow the ordinary machinery actually works

Software

Why your update arrives a week after somebody else's

Releases are handed out in widening circles while somebody watches the failure rate, and your device's position in that queue is not random.

Detailed view of computer code highlighting syntax in colors on a screen.
Photograph by Godfrey Atima via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What follows is the working version of staged rollouts: the decisions in the order you actually meet them, with the reasoning attached.

Before you start

  • Updates go to small cohorts first so faults surface before wide release.
  • Feature switches let code ship disabled and be enabled separately.
  • Firmware and system updates are far harder to withdraw than app updates.

Shipping to everyone at once is untestable

No test environment contains the combination of hardware, settings, languages and accumulated state that real devices present. Faults that appear on a small fraction of devices are invisible in testing and immediately obvious across a large population.

Releasing gradually turns that population into the test, with the crucial condition that the affected group stays small. The alternative is discovering a serious fault after it has already reached everyone, at which point recovery is enormously harder. This is a risk management decision rather than a technical limitation, and it is why the queue exists at all.

Widening circles rather than a single release

Releases usually move through internal users, then volunteers who opted in, then a small percentage of ordinary devices. Each stage runs for long enough to gather meaningful signal, which means days rather than hours for anything subtle. The percentage is increased only if the measurements from the previous stage look acceptable against the previous version.

In the datasheet, devices are assigned to cohorts by an identifier, so the same device tends to sit consistently early or late. That consistency is why some people always seem to receive updates first while others always wait.

What is actually being watched

The most useful signal is the crash rate compared against the version being replaced, measured per model rather than overall. Battery drain, connectivity failures and rates of users reverting a change are watched alongside it as slower indicators. Some problems only appear after a full day of normal use, which is why stages cannot be compressed indefinitely.

The short version: a rise confined to one hardware model or one region will pause the rollout for that group without stopping the rest. Because the comparison is against a baseline, a release during an unusual week can produce misleading measurements.

Shipping code and enabling it are separate acts

Much new functionality ships disabled and is switched on later by a setting the application fetches when it starts. This decouples the slow process of distributing code from the fast process of deciding who should see a feature. It also means two people running the identical version can have genuinely different features available to them.

In the datasheet, switching a feature off remotely is nearly instant, which makes it the fastest available response to a serious problem.

The cost is complexity, since every combination of switches is a configuration that somebody must reason about.

Some updates cannot be taken back

An application can usually be replaced by a corrected version within hours once the fault is understood. System updates that change on-disk formats or firmware are much harder, because reversing them may not be possible safely.

In practice, this is why device firmware moves through the stages more slowly and more cautiously than ordinary applications do. Updates that fix an actively exploited security flaw break the pattern deliberately and are pushed as fast as possible. That urgency is the reason a security update sometimes arrives well ahead of a feature release announced much earlier.

Why your device sits where it does

Hardware model, operating system version, region and language all narrow which build applies to a given device. Devices sold through a network operator often receive a build the operator has approved separately, which adds delay.

Some updates are held for devices in states known to be problematic, such as very low free storage. Manually checking for updates can move a device into an earlier group on some platforms, and does nothing on others. Waiting is usually the sensible option, since the later cohorts receive a version that has already been corrected.

The takeaway

The rollout is an experiment, and the cohort you land in decides whether you are a subject.

Understanding the failure mode tells you more than the feature list does.

Questions readers ask

Is it safer to wait before updating?

For feature releases, often yes. For security fixes the calculation reverses, because the flaw is known and being exploited while you wait.

Why did a feature disappear after an update?

It was probably switched off remotely rather than removed. Feature switches change behaviour without any change to the installed version.

Softwareupdatessoftwarereleasereliability
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida