Docs Audio
Synchronization
Synchronization does not mean "no audible echo". Long before the offset between two speakers grows large enough to hear as an echo, the sound itself starts to break down. This page covers how much offset is acceptable, where it comes from and how it is closed.
What two milliseconds means#
The speed of sound is the distance a sound wave covers in an elastic medium per unit of time. In air, at sea level and around twenty degrees, sound travels roughly 343 metres per second. At that speed a two-millisecond offset corresponds to about 68 centimetres.
This speed does not depend on frequency. Bass and treble travel at the same rate. Were it otherwise, an orchestra would reach the listener in pieces.
What varies is the medium. As the temperature and density of the air change, so does the speed, and sound slows down as the air gets colder. Crossing from warm air into cold air also bends the direction the wave travels in.
There is a reason 343 is used in the calculation. This system runs in climate-controlled indoor spaces, where temperature is held within a narrow band. In shopping centres and comparable venues the indoor conditions follow ASHRAE 55 and TS EN ISO 7730: twenty to twenty-two degrees in winter, twenty-three to twenty-five in summer, with relative humidity between thirty and sixty percent. Because the space stays around twenty degrees all year, a single constant is a reasonable basis.
One of the two speakers behaves as though it were 68 centimetres further away than the other. Across separate rooms that difference disappears. Between two devices standing side by side it is obvious.
What you hear#
An offset of this size is not heard as an echo but as a change in the texture of the sound. The usual descriptions run like this.
- The sound seems wider but somehow odd
- The centre of the vocal drifts; the source does not sit in one place
- There is combing in the upper frequencies
- A faint flanging movement appears
- The two speakers do not focus into one source
The cause is measurable. When two copies of the same signal sum with a small offset, some frequencies reinforce and others cancel. The result is a comb filter, and the position of its nulls comes straight from the delay.
When two copies of the same signal sum with one delayed by Δt, the frequency response of the system is this.
Its magnitude depends only on the delay and the frequency.
Wherever the cosine goes to zero there is a full cancellation. That is where the nulls come from.
For Δt = 2 ms the first null sits at 250 Hz and the rest follow every 500 Hz: 250, 750, 1250, 1750 Hz and onward.
The same delay produces a different phase shift at every frequency, and that shift grows linearly with frequency. This is why the upper frequencies suffer most.
The first null lands at 250 Hz, in the busiest part of the music. Vocals, snares, hi-hats and transients in general are the material this hurts most.
Targets depend on placement#
One number is not correct for every installation. What decides it is where the two speakers stand relative to the listener.
| Placement | Target | Note |
|---|---|---|
| Side by side, or aimed at the same spot | 0.1 - 0.5 ms | Phase relationship must hold |
| Same area, different positions | 0.5 - 2 ms | Acceptable but audible |
| Separate rooms or distant zones | 5 - 20 ms | Tolerable |
The conclusion is that for devices playing side by side the target has to be below one millisecond. In practice a band of ±0.25 to ±0.5 ms is the safe range that keeps the phase relationship intact.
Where delay accumulates#
A shared clock makes every receiver play packets against the same timeline. That alone is not enough, because several links sit between the moment a packet starts playing and the moment the sound reaches the ear.
The total latency of an endpoint is the sum of these links. The last term is the distance the sound travels through the air.
Roughly in order of impact, largest first.
- Sound card and converter latency. The item that varies most between devices.
- Audio driver and buffer settings. Varies with configuration even on the same card.
- The operating system's audio stack. Windows, Linux and small boards each behave differently.
- CPU load and scheduling jitter. Deviation grows on a loaded machine.
- Network adapter. On cable, usually one of the smallest shares.
- Physical distance and room reflections. These come from the installation itself.
The brand of network adapter can add jitter in the microsecond to low millisecond range, and a correct buffer absorbs most of it. The sound card is what actually decides the outcome.
Hardware differences#
Two devices given the "play" command at the same instant can put sound into the air at different moments. The figures below are typical values given to show orders of magnitude, not measurements.
| Output | Typical latency |
|---|---|
| Analogue output of a single-board computer | ~18 ms |
| USB audio converter | ~31 ms |
| Windows shared-mode output | ~60 ms |
| HDMI audio | ~80 ms |
| Bluetooth speaker | 150 ms and above |
Left uncalibrated, these differences are audible even when the network side is flawless. Bluetooth should not be used for this; its latency is both large and unstable.
Endpoint profile#
On mixed hardware, the way to hold synchronization is for every endpoint to carry its own latency profile.
device_id
audio_backend
output_device
measured_output_latency_ms
manual_offset_ms
calibrated_offset_ms
Reading that profile, the system plays some devices deliberately early and others late. The goal is not that the command is issued at the same instant, but that the sound meets in the air at the same instant.
The correction is computed against the slowest endpoint. Each device waits out the difference between its own latency and the largest one.
Endpoint A → 0.0 ms
Endpoint B → -3.4 ms
Endpoint C → +1.8 ms
Network sync and acoustic sync#
These are two separate ideas, and confusing them is a common mistake.
Network sync means every receiver shares one timeline and plays the same packet against the same timestamp. NOT Cable does this. The shared clock flows independently of audio, so a track starting hours later still starts in sync.
Acoustic sync means the sound reaches the ear at the same instant. That requires accounting for the rest of the chain as well: converter latency, buffering, speaker processing, distance.
The second is the correct target. Playing packets at the same instant is necessary, but not sufficient, and that is precisely why NOT Cable builds its network sync so tightly. Acoustic sync can only be built on top of a solid network sync. When the timeline is not consistent below 1 ms, correcting the rest of the chain is meaningless, because every correction is made on shifting ground.
On top of that ground, NOT Cable compensates the rest of the chain per receiver. Every receiver has its own output latency. The buffer depth reported by the sound card driver is accounted for automatically. The remaining difference, from the converter, speaker processing and distance, is corrected with a per-receiver timing offset at 0.1 ms resolution. Two speakers in the same room reach the ear at the same instant even when they run different buffers, and a speaker on the far wall can be set to play as much earlier as its distance requires.
Network sync is what the system gives you. Acoustic sync is the result tuned per receiver on top of that ground. NOT Cable keeps the two apart because they are two different problems, and the second can only be solved once the first is.
Practical advice#
The easiest path is keeping the hardware identical wherever synchronization matters.
- The same board model and the same audio output
- The same sample rate and the same buffer setting
- The same amplifier and speaker type
- A wired network connection where possible
Under those conditions the difference between devices stays small on its own. Mixed hardware works too, but then per-endpoint latency measurement and correction becomes unavoidable.
For network requirements see Network requirements, and for how the clock is carried see NMC, NAP and .nfa.
Prepared by: Hakan E.