The window the square-root law has
Assumes: The bandwidth noise sees · The frequency a sample rate invents · The floor a resistor sets
The filter an average is established the quantity everything below rests on: a mean taken over a window of length is a convolution with a rectangle, and the area under its squared response is exactly . The arithmetic that follows is what every instrument’s specification uses — a reading’s noise is a density times , and averaging readings divides the noise power by — and both statements are about a continuous mean.
Almost no mean is continuous. A voltmeter sums samples, a converter’s output is decimated, a
microcontroller runs y ← y + a(x − y) on a timer interrupt. Between the density and the average sits
a sampler, and a sampler does two things to a noise that neither of the statements above contains: it
folds down everything above half its rate, and it delivers samples that are correlated unless the
filter in front of it is wide enough for them not to be.
Both of those turn out to be one piece of arithmetic, and it has a closed form — which is the shape the habit here is, because a residual with a closed form is a measurement and one without is a disappointment. What a network answers is where that discipline is set out.
Two failures, and why they are one sum
Put a one-pole filter of time constant in front of the sampler — a resistor and a capacitor, or the recursion, whichever the instrument is — and sample it at . The autocorrelation of what comes out is not an approximation:
because a one-pole’s autocorrelation decays exponentially in time and the samples are that exponential read at equal spacings. So the variance of a mean of of them is a double sum that collapses:
with everything in it exact and the sum itself geometric twice over. That single expression carries both departures. When is small the bracket is and the variance is , which is the square-root law. When is not small the bracket is larger and the mean is worse than , because consecutive samples carry some of the same noise.
Where the folding lives is the subtler half, and it is not in that bracket at all. It is in . One sample carries the whole of what the front end passed, which is the density times the front end’s own noise bandwidth — and if that bandwidth runs past half the sample rate, the noise above half the rate is still in every sample. It has not been attenuated. It arrives at full amplitude, indistinguishable from content that was never there, and a mean of the samples averages it exactly as faithfully as it averages the signal.
So there are two claims to keep apart, and the essay’s whole difficulty is that they are usually made in the same breath.
“Averaging N readings divides the power by N” is a claim about samples. It fails only through correlation, so it fails at the narrow end and is exact at the wide one.
“The reading’s noise is the density over twice the window” is a claim about a density. It carries the first claim and the front end’s bandwidth, so it fails at both ends — and the second failure is the one nobody writes down.
One curve, two slopes and a crossing
Divide the exact variance by what the rule gives — a density times — and the result is a function of one variable: the front end’s noise bandwidth measured in half-sample-rates, .
| the rule is out by | the law is out by | ||
|---|---|---|---|
| 0.5 | 0.368 | ×1.080 | ×2.161 |
| 1 | 0.135 | ×1.312 | ×1.312 |
| 2 | 0.018 | ×2.075 | ×1.037 |
| 5 | ×5.000 | ×1.00009 | |
| 10 | ×10.000 | ×1.000000 | |
| 20 | ×20.000 | ×1.000000 |
Read the third column downwards. Past it is simply : the reading’s noise power is the front end’s noise bandwidth divided by half the sample rate, times what the rule says. A front end with 75 kHz of noise bandwidth sampled at 30 kHz gives a reading five times noisier in power than the arithmetic predicts, which is 2.24 times in volts, and nothing about the reading says so — the averaging worked perfectly, and the fourth column proves it did.
The fourth column is the finding that surprises. At the wide end the square-root law is exact to a part in . The folding is not a failure of averaging; it is a failure of the number the averaging was applied to. So a designer who verifies the behaviour on the bench — average four times as long, watch the noise halve, conclude the chain is understood — has verified the one statement that was never in doubt.
Both limits have closed forms with no in them, which is what says the folding belongs to the sampler rather than to the window. With ,
and the second of those is the first multiplied by , because is . Which of the two expressions belongs to which quantity is not a detail: the check as first written held the second against the law and was out by a factor of five, in the direction that would have made the folding look like a defect in the averaging.
The window, and how narrow it is
The curve is monotone, so it crosses the rule exactly once. That is a sharper statement than it sounds: the arithmetic every instrument’s noise specification is written in is right at one front-end bandwidth and wrong everywhere else, and the bandwidth is not half the sample rate.
For a mean of six hundred samples it is 0.1355 of half the sample rate. Around it, the band inside which the rule is right to one per cent runs from 0.0712 to 0.2053 — a factor of 2.88.
| samples in the mean | exact at | right to 1% from | to | a window of | so |
|---|---|---|---|---|---|
| 60 | 0.290 | 0.256 | 0.324 | ×1.27 | 6.17 |
| 200 | 0.195 | 0.146 | 0.245 | ×1.69 | 8.15 |
| 600 | 0.136 | 0.0712 | 0.205 | ×2.88 | 9.74 |
| 2 000 | 0.0908 | 0.0245 | 0.185 | ×7.54 | 10.83 |
| 6 000 | 0.0630 | 0.00831 | 0.177 | ×21.3 | 11.28 |
The last column is the one to take away and it is the answer to the question the essay before it left open. The upper edge of the window has no in it in the limit — it is where reaches 1.01, at — so the sample rate must exceed the front end’s noise bandwidth by 11.53 for the reading’s noise to be within one per cent of the rule. Not by two. The sampling theorem’s factor of two is a statement about recovering a signal without ambiguity, and it has never been a statement about noise; a front end sitting exactly at half the sample rate leaves the reading 28 per cent high in power, which is 1.31 decibels and is the size of error that gets attributed to a part’s data sheet.
The lower edge moves the other way, and it is the front end becoming narrower than the reading’s own bandwidth. In those units the requirement is that exceed by about fifty, and at sixty samples the two edges have nearly met — the window is 1.27 wide, so there is barely a front-end bandwidth at which the rule is right at all. A short average is the case where the rule has no range, which is the opposite of the intuition that a short average is the simple case.
A worked chain, and where the decibel and a third goes
The tables are ratios, and a ratio is easy to nod at. One chain with numbers in it is harder.
A sensor’s output is filtered by a single resistor and capacitor at 15 kHz, sampled at 30 kHz by a converter, and six hundred samples are averaged to make one reading every twenty milliseconds. Every choice there is the obvious one: the filter is at half the sample rate, which is what the sampling theorem is usually read as asking for, and the window is a whole number of mains cycles, which is what the filter an average is shows a mean’s nulls are worth.
The arithmetic a specification would use: the window is 20 ms, so its noise bandwidth is 25 Hz, and the reading’s noise is the density times the square root of 25. With a density of 20 nanovolts per root hertz that is 100 nV.
What the chain delivers is 114 nV. The one-pole at 15 kHz has a noise bandwidth of π/2 times 15 kHz, which is 23.6 kHz — 1.57 times half the sample rate — so the ratio is 1.57 and the reading’s noise power is 1.30 times the rule’s. A decibel and a third, on a measurement whose whole purpose is the last few nanovolts, arriving from a filter placed exactly where the textbook says to place it. That π/2 is the bandwidth noise sees measured on a solved resistor and capacitor, and it is the factor the whole chain turns on.
Putting the filter at 2 kHz instead — a ratio of 0.21, just inside the window — brings the reading to 100.6 nV and costs nothing else whatever, because the signal was never near 2 kHz. Putting it at 60 kHz, which is what happens when somebody chooses a part four times faster than they need and does not recompute, gives a ratio of 6.28 and a reading of 251 nV: two and a half times the specification, with the averaging working perfectly and the square-root check passing.
That is the practical shape of the whole essay. The front end’s bandwidth is a noise parameter of the averager, not only of the signal path, and the two are chosen by different people at different times.
The other end, where averaging longer buys nothing
Below the window the curve falls, and a reading quieter than the rule predicts is not a windfall.
It means the front end has already narrowed the noise below what the window would have, and the correlation the narrowing introduced is what stops the mean improving. At the square-root law is out by a factor of 2.16 in power: a mean of six hundred samples is a mean of two hundred and seventy-eight independent things. Make the mean twice as long and the variance falls by two, as it should, but the reading was never as good as six hundred samples of anything.
Taken to its limit this is the observation instruments are actually designed around. When the front end is very much narrower than the reading’s bandwidth, every sample in the window is nearly the same sample, the mean is a mean of one thing, and the reading’s noise is the front end’s own and does not improve with at all. The chain’s noise bandwidth is the narrower of the two, which is obvious stated that way and is exactly what the rule’s arithmetic cannot express, because contains no front end.
So the sensible reading of the whole curve is short. A chain’s noise bandwidth is a cascade, and is one factor of it. The window measured here is the band over which the other factors happen to cancel — the front end removing as much as the correlation gives back, and the folding not yet arrived — and calling that band “the rule” is what makes both of its edges invisible.
The same variance, from samples
Every number above came from a correlation sum, and a closed form nobody checked is a closed form. The second route shares only the definition of the filter.
A seeded white sequence is run through the recursion with — the same one-pole, as a difference equation rather than as an autocorrelation — normalised so that one sample has unit variance, and six hundred readings of six hundred samples each are taken. The variance of those readings comes out at 0.95 of the closed form’s, against a standard error of 0.058 on an estimate from that many readings — the same square root of two over the count that the average a square root pulls low is about, applied here to the estimate rather than to the reading. That is 0.8 of a standard error, which is agreement, and it is the loosest check in this essay by a wide margin because a variance estimated from readings improves only as .
The correlation sum itself has two routes as well, and it needed them. Written the first time it dropped a term of the second geometric series and was out by 1.7 parts in a thousand at every — small enough to read as quadrature error in a figure and constant enough to prove it was not one. The closed form and the same sum taken term by term now agree to a part in , which is what makes the tables above measurements rather than evaluations.
The folded noise is not white in the reading
One simplification above is worth unpicking, because it is where the folding stops being a factor and becomes a shape.
Everything so far multiplies a flat density by a bandwidth. Under that assumption the folding is a single number: the noise from every band above half the sample rate lands inside the reading’s bandwidth, and because the density is the same everywhere it adds up to one factor. The image bands are interchangeable, so only their total matters.
They stop being interchangeable the moment the density is not flat. What reaches the reading is the sum of the density times the front end’s squared response at every image frequency, and if the density rises, falls or has a line in it, that sum is a different function of frequency from the density itself. Two consequences, and the second is the one that puts a number on a bench.
A rising density folds down heavily. A front end whose own noise rises with frequency — an amplifier’s current noise through a source impedance that rises, which the floor a circuit has measures — has more density in its image bands than in its baseband, so the folded contribution is larger than a flat density’s and the correction above is optimistic.
And a line above half the rate lands at one frequency. A switching converter’s fundamental at 440 kHz, sampled at 30 kHz, arrives at 440 − 14×30 = 20 Hz. It is not noise in the reading; it is a twenty-hertz sinusoid, at whatever amplitude the front end left it, and a mean over 20 ms rejects it by 14.4 decibels rather than by anything about its 440 kHz. This is the failure that reads as drift: a slow wander of a reading with no slow mechanism anywhere in the circuit, produced entirely by something fifteen times faster than the sampler. It is the same shape the millivolts in the wire describes at direct current, with a sampler in the middle moving the frequency.
So the window measured in this essay is the right statement for a flat density and a floor for anything else. Where the density is not flat the upper edge is nearer than the table says, and where it carries a line there is no bandwidth at which the rule applies at all — the reading has a tone in it, and a noise bandwidth is the wrong quantity to be computing.
What is left out
The front end is one pole. A real anti-alias filter is not, and a steeper one changes both edges: its noise bandwidth is a smaller multiple of its corner, so it folds less for a given passband, and it correlates samples differently because its autocorrelation is no longer a single exponential. The closed sum above is specific to the one-pole and the shape of the finding is not — a window with two edges, whose upper one is about folding and whose lower one is about correlation. The ratio that does not walk to one is where the noise bandwidths of the realised families are measured, and each of them would put the upper edge somewhere else.
The density is flat. Everything here multiplies a density by a bandwidth, which is right while the density does not vary across it. On a flicker density it does, and the corner where averaging stops working is what that costs: a longer window narrows the bandwidth and the density under it rises to meet it, so the lower edge of this essay’s window is not the only reason averaging longer stops helping.
And the samples are of noise and nothing else. A real reading has a signal in it, the front end band-limits that too, and the settling of the front end after a change is a separate quantity from any of these — the filter an average is measures it for the three averagers and finds a factor of seven between them at one noise bandwidth. A chain chosen for the folding requirement here has a settling time that follows from it and is not free.
Still open: the filter that makes both edges move, and the decimation nobody counts
The order of the front end, priced in sample rate. The upper edge is 11.53 for a one-pole, and it is that large because a one-pole’s noise bandwidth is π/2 times its corner and its skirt admits noise from decades above. A steeper front end folds less, so the required ratio of sample rate to corner falls — and the same sweep over the realised families would say by how much, which converts an order into a clock rate exactly as what the filter in front costs converts one into an oversampling ratio for a signal.
Decimation in stages. Nothing here decimates. A real chain samples fast, filters digitally, discards samples, and repeats — and every stage of that is a sampler with a front end, so every stage has the window measured here. Whether the windows compose, or whether an early stage’s folding is irrecoverable by a later one’s filter, is a question the same correlation sum would answer over a cascade rather than over one mean.
The measurement that would find it on a bench. The two failures move in opposite directions with the front end’s bandwidth, so they are separable by a sweep that nobody performs: vary the sample rate with the front end fixed and the folding moves while the correlation does not. A plot of a reading’s noise against sample rate at fixed window length should be flat, and where it is not is the whole of this essay in one instrument.
What is checked
The correlation sum is computed twice. A closed form in and , and the same sum taken term by term, required to agree to a part in . They now agree to ; the first version of the closed form agreed to 1.7 parts in a thousand, at every , which is the signature of a dropped term rather than of an approximation.
Each limit is held against the quantity it belongs to. The rule’s error against and the square-root law’s against — with the tolerance being the finite- term itself, , rather than a fixed number. A flat failed at , which is exactly that quantity, so the bound is now the size of the approximation being made rather than a figure chosen to pass.
That the square-root law is exact at the wide end is required separately, because the essay’s central claim is that the folding is not a failure of averaging and a reader shown only the first curve would conclude that it is.
The curve is required to be monotone over seven decades of front-end bandwidth before either edge is bisected. Both edges are found by bisection on a monotone function, and a bisection on a function that is not monotone returns an end of its own bracket — which is how the lower edge was first reported, as , the bottom of the range it was searched over.
And the variance is measured a third way, from seeded samples, at 0.8 of a standard error. It is the loosest agreement in the essay and it is the only one that shares no arithmetic with the rest.
Part 6 on noise bandwidth
One argument about Noise bandwidth, and one of 6 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the idea: the things themselves, and every essay that touches each one.
AliasingAveragingEquivalent noise bandwidthModel rangeSampled dataSeeded generatorSpectral densityVerification
- A resistor made of a clock aliasing, model range, sampled data
- Only the real part is warm model range, spectral density, verification
- The amplifier inside the sample aliasing, equivalent noise bandwidth, spectral density
- The bandwidth a bin is not equivalent noise bandwidth, spectral density, verification
- The cross terms that outnumber the sources seeded generator, spectral density, verification
- The noise a true-RMS meter reads low averaging, model range, seeded generator