The question this page answers
Coherent averaging only works because every captured block is rotated back onto a common phase before it is summed - the per-block de-rotation the FFT analyser chapter describes. That rotation angle is itself measured, so it carries a small error. The natural worry: if the de-rotation angle is slightly wrong on every block, does that error quietly eat into the averaged harmonic levels - planting a false floor, or making the harmonics read low on a long run?
This page works the numbers. The short answer is reassuring: the magnitude accuracy of a coherent average is set by the bin signal-to-noise ratio and the number of averages - the ordinary 1/√n law - not by the de-rotation. The de-rotation's own footprint on the measured levels is microscopic (micro-decibels) and, crucially, it does not grow with averaging time, provided the rotation angle is measured afresh on every block rather than integrated from block to block. The rest of this page shows why, and what it costs to get it wrong.
The worked example
All the numbers below are for one concrete, demanding case: a faint tone averaged very deeply.
| Quantity | Symbol | Value |
|---|---|---|
| Sample rate | fs | 384 000 Hz |
| FFT length | N | 2²¹ = 2 097 152 |
| Block duration | T = N/fs | 5.46 s |
| Bin spacing | Δf = fs/N | 0.183 Hz |
| Tone | f0 | 1000 Hz, bin-snapped |
| Residual drift | δf | ±0.06 ppm ≈ ±60 µHz |
| Tone level | - | −80 dBV |
| Per-bin noise floor | - | −145 dBV (per bin) |
The quantity that decides everything is each line's per-bin signal-to-noise ratio ρ - the tone or harmonic energy in its bin against the noise energy in that same bin. For this example the fundamental and the first two measurable harmonics sit at:
| Line | Level | Per-bin SNR ρ |
|---|---|---|
| Fundamental | −80 dBV | 3.2×10⁶ (65 dB) |
| 2nd harmonic | −130 dBV | 31.6 (15 dB) |
| 3rd harmonic | −135 dBV | 10 (10 dB) |
Higher harmonics here sit below the per-bin floor and only emerge once averaging has lowered it; the de-rotation reference is built from the lines that are above it.
Why the angle has to be re-measured every block
With the tone snapped exactly onto a bin and blocks read contiguously, the block-to-block phase step would be zero. The few-ppm difference between the DAC and ADC clocks leaves a tiny residual: the tone creeps by a fraction of a bin, which shows up as a slow phase walk from one block to the next of
per block within a contiguous run. The within-block frequency offset is a few ten-thousandths of a bin, so the level lost to scalloping is parts in 10⁷ - negligible. The drift is therefore not a magnitude problem; it is a walking phase that the de-rotation strips block by block. The regime is noise-limited, not drift-limited - as long as the rotation angle for each block is measured on that block, not predicted from the last one (the reason becomes concrete in the non-contiguous-frames section).
How accurately the de-rotation angle is known
In one block the fundamental bin holds the tone phasor plus a noise phasor. The angle read off it is the true phase plus an error whose size, at high SNR, is the perpendicular noise component divided by the signal amplitude:
So a single block fixes the de-rotation angle to about two-hundredths of a degree[1]. That is the precision of the whole method - and it comes almost entirely from the 65 dB fundamental.
One could instead pool the fundamental and the harmonics: each harmonic h rotates h× as fast under a timing slip, so each gives an independent reading of the common drift. Combined as a weighted least-squares fit, the angle variance becomes
but with a 65 dB fundamental, Σ h²ρh = 3.2×10⁶ + 216 - the harmonics add 0.007 %. Pooling buys essentially nothing here; both methods give σθ ≈ 0.023°, fundamental-dominated.
Does the error build up over a long average?
This is the heart of the matter. The per-block angle error is independent and zero-mean from block to block. Two very different things can happen with such an error, depending entirely on how the reference is carried:
- Measured absolutely each block (or held by a closed tracking loop): the errors are independent and zero-mean, so they average. The applied per-block correction keeps its 0.023° spread, but the phase of the accumulated tone tightens as
- Integrated from block to block by a free-running counter: the same small errors now accumulate - their variance grows in proportion to the number of blocks, the reference wanders off, and the tone slowly de-coheres (its level sags, worse the longer you average). This is the one mechanism that would plant a magnitude error that grows with time.
The rule that follows is absolute: keep the de-rotation reference measured-per-block or held by a bounded loop; never let it free-run. Phonalyser does exactly this, which is what makes the 8-12-hour averages in the FFT chapter possible.
What it does to the measured levels
After de-rotation, harmonic h carries a residual rotation of −h·(block error). Accumulating and expanding that small rotation splits its effect into two cleanly separable parts.
A constant magnitude bias - micro-decibels
The averaged amplitude of a randomly-jittered phasor sits a hair below the true amplitude:
For this example that is −0.69 µdB at the fundamental, −2.7 µdB at the 2nd harmonic, −6.2 µdB at the 3rd. It is the only genuinely de-rotation-induced level term, it is constant (averaging cannot remove it) - and at a few micro-decibels it is irrelevant.
A random level error - the ordinary 1/√n
The additive noise in each bin also moves the measured amplitude, and this is what actually limits magnitude accuracy:
Putting the two side by side over a range of averaging depths shows the gap (1σ, in dB):
| averages n | F - de-rot bias / random | H2 random | H3 random |
|---|---|---|---|
| 1 | −0.7 µdB / 3.5×10⁻³ | 1.09 | 1.92 |
| 100 | −0.7 µdB / 3.5×10⁻⁴ | 0.109 | 0.194 |
| 1000 | −0.7 µdB / 1.1×10⁻⁴ | 0.034 | 0.061 |
| 10000 | −0.7 µdB / 3.5×10⁻⁵ | 0.011 | 0.019 |
For every line the de-rotation bias is four to six orders of magnitude below the additive-noise random error. Meanwhile the noise floor itself falls by 10·log₁₀n - to −165/−175/−185 dBV at n = 10²/10³/10⁴ - which is what lifts the buried harmonics into view. The de-rotation bias plants no floor that would block that. Magnitude accuracy is set by bin SNR and n, full stop.
Non-contiguous blocks after glitch rejection
The discontinuity detector throws out corrupted blocks - that filtering is precisely what lets a long average converge at all, because the accumulator only ever sees clean blocks. But the survivors are not contiguous: arbitrary gaps sit where rejected blocks were dropped.
Across such a gap the tone's absolute phase has advanced by an unknown whole-plus-fraction number of cycles, so the angle the next clean block needs is effectively uniform over the full ±180° - not the 0.12°-per-block creep of a contiguous run. The de-rotation must cover the whole circle. This is harmless on one condition: the angle is measured afresh on each block (read directly from that block's fundamental), never predicted from the previous block or carried by a free-running counter - which cannot know how long the gap was. This is the same caveat as before, here promoted from advisable to mandatory: non-contiguous blocks force absolute de-rotation.
It costs nothing in coherence. The bin model is identical for any absolute phase; de-rotation removes that phase exactly; only the estimation error propagates, and its size - equation (2) - does not depend on the value of the absolute angle. However wild the angles are, the analysis above holds verbatim: the arbitrary absolute phase washes out, the residual is pure noise. There is no harmonic wrap-ambiguity either, because rotating bin h by h×(the measured fundamental angle) gives the correct rotation for any amount of phase wrap.
Locking to the running average - without self-poisoning
How the fundamental de-rotation is actually formed:
- An absolute model, anchored to the first block. Each block is rotated by an angle proportional to its absolute sample distance from the first contributing block, using a sub-bin frequency κ pinned at that first block (harmonics ride h× that angle). This is the absolute, never-integrated rotation the sections above require.
- The frequency refined once, over the first 24 clean blocks. The de-rotated fundamental's slow phase slope is fitted across those blocks - a baseline far longer than one block hop - and folded into κ a single time. That fit is about a hundred times tighter than a single-block peak estimate, and it kills the slow ramp that a residual κ error would otherwise cause on a long average (the "why κ is fitted over the whole segment" point).
- Then a gentle phase lock to the deep average. Each new block's freshly-read fundamental phase is compared with the phase of the accumulated tone (the √m-stabilised reference, not the first block), and a small fraction - a tracking gain of 0.10 - of the difference is folded back into the running time origin. Full correction is applied only as a one-shot realignment after a genuine re-sync jump.
Does locking to a noisy average poison it?
The accumulated reference is itself an estimate - its phase variance is σθ²/m after m blocks. Could the loop fold that noise into a magnitude error? No, and the bound is small and shrinking:
- The reference stiffens, it does not drift. A new block moves the accumulated angle by only ~1/m, so the reference locks to an ever-more-stable consensus.
- The level footprint vanishes with depth. Averaged over a run of depth M, the reference-noise contribution to a harmonic's level works out to
- - nano-decibels at the fundamental (3×10⁻⁸ / 5×10⁻⁹ / 6×10⁻¹⁰ dB at M = 10²/10³/10⁴), and getting smaller with depth, six orders below the additive-noise random error.
- No walk-away. The loop is driven to zero in the difference between the fresh block and the average, but free in the common mode - the whole accumulated phasor may rotate slowly as one. A common rotation leaves every level and every relative harmonic phase intact; only the reported absolute phase wanders, which does not matter for levels or THD.
- The low gain is the guard. At full gain the loop would fold each block's whole noise into the time origin, which would random-walk and de-cohere early-versus-late blocks. The 0.10 gain low-passes it to a steady-state jitter well below σθ; full gain fires only on a real re-sync jump, where the difference is signal, not noise.
So the running-average reference carries noise, but its level effect vanishes as (ln M)/M rather than accumulating - the total bias stays the constant few-micro-dB of equation (5).
Overlap: the effective number of averages
The 1/√n law assumes independent blocks. Overlapping blocks share samples, so their bin noise is correlated and the variance does not fall as fast. The correct substitution is an effective independent count:
where ρj is the window's self-overlap at hop j, non-zero only while blocks still overlap (4 neighbours at 75 % overlap, 8 at 87.5 %, 16 at 93.75 %). Every random-error term above takes n -> neff.
The striking result concerns the floor reached in a given amount of wall-clock time. The number of blocks in time T is n = T/((1−ov)·N), so the achievable noise-power floor scales as NENBW·(1−ov)·Fcorr - and in the high-overlap limit this product collapses to a window-independent identity:
Computed for four very different windows, the floor-per-unit-time relative to the best:
| window | NENBW (bins) | 75 % | 87.5 % | 93.75 % |
|---|---|---|---|---|
| Blackman-Harris (−92 dB) | 2.00 | 1.00 | 1.00 | 1.00 |
| Flat-top (5-term) | 3.77 | 1.00 | 1.00 | 1.00 |
| Dolph-Chebyshev (−220 dB) | 2.86 | 1.00 | 1.00 | 1.00 |
| HFT248D (−248 dB) | 5.65 | 1.22 | 1.00 | 1.00 |
Two consequences fall straight out, and they match Heinzel's recommended overlaps independently[2]:
- Above each window's recommended overlap, all four windows reach the same floor in the same time. A wider main lobe (more noise per block) is repaid exactly by finer decorrelated hopping. The wide HFT248D only catches up by 87.5 %; below that it pays a penalty.
- 93.75 % buys no noise benefit over 87.5 % - the floor is identical - it merely doubles the transform rate. Above the recommended overlap, raising it further is pure compute waste.
Reaching neff = 100 / 1000 / 10000 independent-equivalent averages therefore takes roughly 4.5 min / 45 min / 7.6 h of capture at this block length - the deep-average rows are genuinely long runs.
Which window - and why it is not chosen for noise
Because the floor-per-time is equal above the recommended overlap, the window is chosen on the other two axes:
| window | NENBW | side-lobe floor | amplitude flatness | role |
|---|---|---|---|---|
| Blackman-Harris (−92 dB) | 2.00 | −92 dB | 0.83 dB | narrowest, cheapest noise |
| Flat-top (5-term) | 3.77 | ≈ −88 dB | ±0.01 dB | amplitude accuracy off-bin |
| Dolph-Chebyshev (−220 dB) | 2.86 | −220 dB | - | deep uniform side-lobes |
| HFT248D (−248 dB) | 5.65 | −248 dB | 0.0007 dB | deepest side-lobes + flat |
- Amplitude flatness is not needed here. The tone is bin-snapped and drifts only a few ten-thousandths of a bin, so even Blackman-Harris's 0.83 dB half-bin scallop costs parts in 10⁷ of a dB. Flat-top's signature advantage is wasted on a bin-locked tone - which makes the plain flat-top, with its wide main lobe and only −88 dB side-lobes, the worst pick for this task.
- Leakage from the −80 dBV fundamental is the real differentiator. For a bin-exact tone a cosine-sum window nulls at every integer offset, so inter-harmonic leakage is near zero. But drift and re-sync glitches push the tone a fraction of a bin off centre, and the floor it then plants at distant bins is set by the side-lobe level: Blackman-Harris's −92 dB puts a −172 dBV pedestal under the spectrum - above the −185 dBV harmonics dug out at n = 10⁴, so it could masquerade as signal. The deep-side-lobe windows put that pedestal at −300 dBV and below - 130 dB of margin that guarantees a dug-out harmonic is the device, not fundamental leakage.
Bottom line
- The de-rotation angle is known to σθ ≈ 0.023° per block (fundamental-dominated). It does not accumulate, and its level footprint is a constant −h²·0.69 µdB.
- Magnitude accuracy is the ordinary coherent 1/√n set by bin SNR - the de-rotation bias is four to six orders below it and plants no floor.
- The reference must be measured-per-block or bounded by a loop, never free-integrated; locking to the deep average is safe because its own noise costs only ½h²σθ²·(ln M)/M - vanishing, not accumulating.
- Overlap enters only through neff = n/Fcorr; above each window's recommended overlap the floor reached per unit time is window- and overlap-independent, so HFT248D at 87.5 % is chosen for side-lobe margin, not noise.
References
- 1. The 1/√(2ρ) phase-estimation error of a tone in noise is the high-SNR Cramér-Rao bound - Wikipedia: Cramér-Rao bound.
- 2. Equivalent noise bandwidth, window self-overlap correlation and the recommended-overlap values - G. Heinzel, A. Rüdiger, R. Schilling, "Spectrum and spectral density estimation by the discrete Fourier transform (DFT), including a comprehensive list of window functions", 2002.
◀ FFT analyser · Theory of operation · next: Frequency response ▶