Skip to content跳到正文
All notes全部笔记
2 min read阅读约 2 分钟

Recording latency in the browser is not a constant浏览器里的录音延迟不是一个常数

Why "let the user drag the waveform into place" is the wrong interaction, and how to align a take automatically by cross-correlating energy envelopes.为什么"让用户手动对齐波形"是个错误的交互,以及用能量包络互相关自动对齐的做法。

Web AudioDSPAudio ProgrammingWeb AudioDSP音频编程

Record vocals against a backing track in a browser and the take comes back sitting tens of milliseconds late. The reasons are not mysterious:

  • Output buffer latency (backing track from AudioContext to your ears)
  • Input buffer latency (microphone to MediaRecorder)
  • The device's own A/D and D/A conversion delay
  • Bluetooth headphones, adding another 100–200 ms

The problem is that the sum is not a constant. It changes with the device, with the sample rate, with system load — two sessions on the same machine can differ by a dozen milliseconds. Hardcoding a compensation value is pointless.

Manual alignment is the wrong answer

The usual approach hands the user a waveform and lets them drag it into place. It looks professional, but it pushes work a machine should do onto a person — and the person does it worse, because eyeballing waveform alignment is nowhere near as accurate as an algorithm.

Cross-correlating energy envelopes

The key step: do not correlate raw samples. A vocal waveform and a drum waveform have nothing in common, so correlating them directly returns noise. But they do share one thing — transients. A sung consonant and a snare hit are both a spike in the energy envelope.

So reduce both signals to RMS energy envelopes first, then scan a physically plausible latency window:

const WINDOW_MS = { min: -40, max: 240 };

function bestOffset(vocalEnv: Float32Array, backingEnv: Float32Array) {
  let best = { offsetMs: 0, score: -Infinity };
  for (let ms = WINDOW_MS.min; ms <= WINDOW_MS.max; ms += 1) {
    const score = dot(vocalEnv, shift(backingEnv, ms));
    if (score > best.score) best = { offsetMs: ms, score };
  }
  return best.offsetMs;
}

A few practical details:

  • Use only the first 20 seconds. It is enough. Further in, a timing offset is more likely to be the singer genuinely pushing or dragging the beat — that is not latency and should not be "corrected" away.
  • Window from -40 ms to +240 ms. Leave headroom on the negative side, because some devices report timestamps that run ahead; the 240 ms upper bound covers the Bluetooth worst case. Any wider and it starts matching the wrong beat.
  • A 1 ms step is plenty. At 48 kHz that is 48 samples, far below the threshold of perception.

The result

Stop the recording, and in the moment it takes to enter the archive window this runs in the background and the waveform snaps into place. The user never needs to know the concept of latency exists.

That is the only part of a tool like this I think is worth building: when an algorithm can determine something, do not turn it into a knob.

在浏览器里对着伴奏录人声,录完你会发现人声整体偏后几十毫秒。原因不神秘:

  • 输出缓冲区延迟(伴奏从 AudioContext 到耳朵)
  • 输入缓冲区延迟(麦克风到 MediaRecorder)
  • 设备本身的 A/D、D/A 转换延迟
  • 蓝牙耳机再加 100–200ms

问题在于这个总和不是常数。换设备变、换采样率变、系统负载高了也变,甚至同一 台机器两次会话都能差十几毫秒。所以硬编码一个补偿值没有意义。

手动对齐是个错误的答案

常见做法是给用户一个波形,让他自己拖着对齐。这个交互看起来很专业,其实是把一个 机器该干的活推给了人——而且人干得更差,因为肉眼看波形对齐远不如算法准。

能量包络互相关

关键的一步是:不要相关原始采样。人声的波形和鼓的波形毫无相似之处,直接互相关 的结果是噪声。但两者共享一个东西——瞬态。唱出来的辅音和一记军鼓,在能量包络上 都是一个尖峰。

所以先把两路信号降维成 RMS 能量包络,再在物理上说得通的延迟窗口里扫描:

const WINDOW_MS = { min: -40, max: 240 };

function bestOffset(vocalEnv: Float32Array, backingEnv: Float32Array) {
  let best = { offsetMs: 0, score: -Infinity };
  for (let ms = WINDOW_MS.min; ms <= WINDOW_MS.max; ms += 1) {
    const score = dot(vocalEnv, shift(backingEnv, ms));
    if (score > best.score) best = { offsetMs: ms, score };
  }
  return best.offsetMs;
}

几个实践上的细节:

  • 只取前 20 秒。够了。而且越往后唱,节奏偏移越可能是真的在抢拍或拖拍,那不 是延迟,不该被"修正"掉。
  • 窗口取 -40ms 到 +240ms。负数一侧留一点余量,因为有些设备的时间戳会跑到 前面去;上界 240ms 覆盖到蓝牙的最坏情况。再宽就开始匹配到错误的拍点上。
  • 步长 1ms 就够。48kHz 下那是 48 个采样,对人耳来说远低于可察觉阈值。

结果

录完点停止,进入归档窗口的那一刻后台跑完这件事,波形自己吸附上去。用户不需要知道 延迟这个概念存在。

这是我觉得这类工具唯一值得做的地方:算法能确定的事,就不要做成一个旋钮。