Tech Behind ThingsHow the ordinary machinery actually works

Networks

Why A Voice Call Sounds Worse Than A Voice Note

Recorded audio can be compressed carefully and resent when it fails, while a live call has milliseconds to decide, and that deadline is what you hear.

Focused detail of a modern server rack with blue LED indicators in a data center.
Photograph by panumas nikhomkhai via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A voice message sent over the same connection as a call sounds clearly better. The difference is not bandwidth but the deadline the encoder is working to.

A call cannot wait for the network

Conversation breaks down once the delay between speaking and being heard exceeds a couple of hundred milliseconds. Everything in the chain must fit inside that budget.

An encoder therefore has a few tens of milliseconds of audio to work with and no opportunity to look ahead at what is coming next.

A recording faces no such constraint. The whole file exists before encoding starts, so the algorithm can analyse it as a whole and allocate detail where it matters.

Lost packets cannot be requested again

Data transfers recover from loss by asking for a retransmission. By the time a replacement arrived, the moment of a live call would have passed.

Live audio therefore accepts loss. Missing fragments are concealed by repeating or interpolating from surrounding audio, which produces the familiar smeared or robotic effect.

A voice note is a file. Every byte is guaranteed to arrive eventually, so nothing needs concealing and the result is exactly what was recorded.

The rate has to be steady

A file can be compressed unevenly, spending many bits on complex passages and few on silence. The average is what matters and the player buffers accordingly.

A call must produce output at a near-constant rate, because a sudden surge would arrive late. The encoder is capped at every instant, not on average.

This forces compromises during exactly the moments that need detail most, such as two people speaking over one another.

Legacy codecs still appear

Traditional telephone audio discards everything outside a narrow band of frequencies, a decision made when copper capacity was scarce and never fully undone.

Wideband codecs exist and sound dramatically better, but both ends and every network in between must support the same one. Any leg that does not forces a fallback.

This is why call quality can vary between two people using identical handsets. The connection negotiated down to what the weakest link could handle.

Transcoding compounds the loss

Audio compressed once, decoded and compressed again loses more than either step alone, because the second encoder is now compressing the first one's artefacts.

Calls crossing between networks and formats can be transcoded several times without anyone involved being aware of it.

An application that keeps the call in one format end to end avoids all of that, which is the main reason internet calls between the same app often sound better than the telephone network.

Questions readers ask

Is a mesh system better than a single powerful router?

Only where coverage is the limitation. One well-placed unit serving a small flat will beat three nodes relaying through each other.

Do more nodes always improve things?

No. Each wireless hop costs airtime, and nodes that hear each other well compete for the same channel. Two good positions beat four poor ones.

Networkswirelesshome networkcoveragenetworking
Grigor Petrov
Hardware writer, Tech Behind Things

Grigor writes about silicon, thermals and the physical limits designers keep bumping into.

Also by Grigor Petrov