Home Tech
How a smart speaker listens for its name without recording the room
A tiny model running locally on a low-power chip watches for one pattern, and only what follows it is sent anywhere.

There is a settled way of talking about wake word detection. It is worth asking how much of it survives contact with the detail.
The argument in brief
- Wake word detection runs locally on a dedicated low-power processor.
- Audio is held in a short rolling buffer and continuously discarded.
- False triggers are the mechanism by which unintended audio gets uploaded.
A small model on a small chip
A dedicated low-power processor runs a compact model trained to recognise one specific phrase in the audio stream. It is deliberately tiny so it can run continuously on very little power, which means it can recognise almost nothing else.
Audio passes through a short rolling buffer of a second or two and is overwritten continuously. Only when the pattern matches does the device open a connection and begin sending audio.
The trade-off is set deliberately
A detector can be tuned to miss the wake word rarely at the cost of triggering falsely more often, or the reverse. Manufacturers lean towards not missing it, because a device that ignores you feels broken while one that wakes occasionally feels harmless.
Mechanically, that choice is why televisions, conversations and songs sometimes trigger devices, and researchers have catalogued many such accidental triggers. Each false trigger uploads a few seconds of whatever was being said, which is the actual privacy exposure rather than continuous recording.
Beamforming picks a voice out of a room
Multiple microphones spaced apart receive the same sound at slightly different times. Combining the signals with calculated delays reinforces sound arriving from one direction and cancels other directions.
The device scans for the direction giving the strongest match and locks onto it, which is how it hears you over music. It also subtracts its own output from the input, which is why it can hear a command while playing loudly.
What happens after the trigger
The buffered audio plus what follows is sent for speech recognition, which usually happens on a server because the models are large. The transcript is then interpreted to identify an intent and parameters, and a response is generated and returned. Some simple commands are handled locally, which is why they work with the internet disconnected and others do not.
On-device recognition has expanded as models have shrunk, and coverage varies by language and by device generation.
Retention is a settings question
Recordings and transcripts have historically been retained to improve recognition, sometimes reviewed by human contractors, which caused significant controversy when it emerged. Most platforms now offer explicit controls over retention, human review and deletion, often not enabled by the privacy-preserving default. Reviewing those settings once, and periodically deleting stored recordings, is the practical action available.
Rules on consent and retention differ substantially by jurisdiction and have driven several of the changes.
Figures here are typical rather than guaranteed — check the spec sheet for your part.
Physical controls are the honest ones
A mute button that electrically disconnects the microphone is a stronger guarantee than any software setting. Devices differ in whether the mute is a hardware disconnection or a software flag, and manufacturers do not always make this clear. A device with a camera should have a physical shutter for the same reason.
For anyone genuinely uncomfortable, not putting the device in the bedroom is more effective than any configuration.
The takeaway
It listens locally for one word and forgets the rest. The exposure is the false triggers, not a permanent recording.
Once you know what it is trading away, the design stops looking arbitrary.
Questions readers ask
Is my smart speaker recording everything?
It is processing everything locally in a short buffer and discarding it. Audio leaves the device after a wake word match, including when that match was a mistake.
Does muting actually work?
It depends on whether the mute physically disconnects the microphone or merely sets a flag. Check the manufacturer documentation, since both designs are sold.





