Tech Behind ThingsHow the ordinary machinery actually works

Devices

How Several Microphones Decide Which Voice To Keep

Phones and headsets use two or more microphones to work out where a sound came from, then suppress everything arriving from the wrong direction.

A classroom setting featuring laptops and desks, capturing a modern educational environment.
Photograph by Adam Sondel via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A single microphone cannot tell a voice from the traffic behind it, because both arrive as one mixed pressure wave. Adding a second microphone changes the problem entirely.

Distance introduces a tiny delay

Sound travels at a finite speed, so a wave reaching two microphones a few centimetres apart arrives at one fractionally before the other.

That difference is measured in microseconds, but audio is sampled tens of thousands of times a second, which is fast enough to detect it reliably.

From the delay, the device can calculate the angle the sound came from. Two microphones give a cone of possibilities, and a third narrows it considerably.

Beamforming sums the signals selectively

If the signals are delayed to compensate for the path difference and then added, sound from the chosen direction reinforces itself while everything else partially cancels.

The effect is a virtual microphone aimed at a region of space. It is created in software and can be steered without moving anything physical.

This is why holding a phone in a different position changes call quality noticeably. The beam is pointed where your mouth is expected to be, not where it is.

Steady noise is modelled and subtracted

Direction alone does not remove a fan or road noise coming from every side. For that, the device builds a statistical model of the background during pauses in speech.

Whatever persists unchanged is treated as noise and removed from the mixture. Speech, which fluctuates rapidly in pitch and level, survives the filter.

The approach fails predictably in a crowded room. Another conversation looks exactly like speech to the model, so it is preserved rather than suppressed.

Wind is a different failure

Wind does not travel to the microphone as sound. It strikes the diaphragm directly, producing an enormous low-frequency disturbance unrelated to anything happening around you.

Because it arrives at each microphone independently rather than as a wave crossing them, the correlation between channels collapses, and the device can detect the condition from that alone.

Detection usually triggers a switch to whichever microphone is least exposed, plus aggressive low-frequency filtering. The call sounds thin, which is preferable to sounding like a storm.

The processing happens before transmission

All of this runs on the device, in the milliseconds between capture and encoding. The far end receives a cleaned signal and cannot recover what was discarded.

That is a deliberate choice, since the network carries far less data than raw audio. Sending everything and sorting it later is not an option a call budget allows.

It also means the algorithm's mistakes are permanent. When a phone decides your voice was background noise, that portion of the sentence is simply gone.

Questions readers ask

Why does brightness jump when I unlock the phone?

The sensor is often only sampled while the screen is on, so the first reading after unlocking replaces a stale value from earlier.

Does automatic brightness save battery?

Usually yes, because most people set a fixed level high enough for the worst case and then leave it there in dim rooms.

Devicesdisplayssensorsperceptionpower
Grigor Petrov
Hardware writer, Tech Behind Things

Grigor writes about silicon, thermals and the physical limits designers keep bumping into.

Also by Grigor Petrov