The floor, which bounds from below

The corner where averaging stops working

Average N samples and the noise falls by √N. That is a statement about independent samples, and flicker noise's samples are not independent — its correlation extends over every time scale, which is what a spectrum with no bottom means. Measured on the same seeded stream filtered and not: the white sequence falls as the −0.510 power of the block length and the pink one as the −0.087 power, so a thousand-sample average buys a factor of 36.9 on one and 2.1 on the other.

Assumes: The floor a circuit has · The floor a resistor sets · The bandwidth noise sees

Everything the noise field has drawn so far has been white: a density flat with frequency, samples independent of one another, and a root-mean-square that falls as the square root of the number of samples averaged. That last property is the one every measurement in every laboratory is built on. Average for longer and the answer gets better; average four times as long and it gets twice as good; and the only cost is time.

Below some frequency it stops being true, and this essay is about measuring where and by how much.

The previous essay took an amplifier’s two noise generators as numbers. They are not numbers. Below a corner that is a property of the device — its material, its surface, its bias current — the density of both rises, and it rises approximately as 1/f. A spectrum that rises as 1/f as far down as anybody has ever measured is a spectrum with no bottom, and a spectrum with no bottom has consequences for averaging that are not intuitive and are not small.

Averaging a white sequence, and averaging a pink onecomputed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.087 ± 0.013, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 2.1 on the second.100m11101001ksamples in the blockspread, relative to one sampledashed: the √N lawwhite: n to the-0.510 ± 0.006pink: n to the-0.087 ± 0.0131024 samples buy36.9× on whiteand2.1× on pink1/f band1 Hz … 10 kHzsolved, then checked — one stream, filtered and notexponent -0.09 against -0.51
Fig. 1 The same seeded white stream, averaged in non-overlapping blocks, before and after the 1/f network. The dashed line is the √N law. The white sequence’s spread falls as n0.510n^{-0.510} across five seeds; the pink one’s falls as n0.087n^{-0.087}, which is very nearly not at all. The slider is the bottom of the 1/f band, and raising it is what gives averaging its power back.

Building a shape as a circuit

Flicker noise is normally introduced as a density that “goes as 1/f”. That is a shape and not a circuit, and every figure on this site is measured on a circuit — so the first job is to build one.

The standard construction is a chain of lag sections with their poles and zeros interleaved. A pole starts the magnitude falling at 20 dB per decade; a zero part of a decade later stops it; the next pole starts it again. Averaged over the band the slope is half of a single pole’s, which is 10 dB per decade of magnitude and therefore 20 dB per decade of power — and a power density falling at 20 dB per decade is exactly 1/f, because a density is the square of a magnitude.

The chain is a netlist like everything else here: each section is a resistor in series, then a resistor and capacitor to ground, which places a zero at 1/(2πRC) and a pole at 1/(2π(R₁+R₂)C), the pole always below the zero.

Nothing in the construction enforces the slope. It falls out of the interleaving, which is what makes it worth measuring rather than asserting: the figure fits a straight line through the solved magnitude and reports −9.989 dB per decade at one pole-zero pair per decade, and — the check that matters — that this is 0.500 times the slope of one pole on its own, measured on a separate RC over the same two decades.

A 1/f density built from 4 interleaved poles and zeros. computed by solving, not by drawing. 4 lag sections over 4 decades, at 1 pole-zero pair per decade. The magnitude fitted over the band falls at -9.989 dB per decade against a target of −10, which squares to a density of 1/f, and the worst departure from that line is 0.261 dB. Nothing in the construction enforces the slope: it is what a chain of interleaved corners does.
Fig. 2 Four lag sections over four decades, with the fitted line drawn from the fit rather than from the design. The staircase is the price of a finite number of sections and it is shown rather than smoothed: 0.261 dB of worst departure here. The slider is how many pole-zero pairs there are per decade.
Averaging a white sequence, and averaging a pink one. computed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.083 ± 0.015, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 2.1 on the second.
Fig. 3 The same measurement with the pink band beginning at half a hertz. The white sequence still averages down with an exponent of −0.510, near the −0.5 that theory gives; the pink one manages −0.083, and a thousand and twenty-four samples buy a factor of 2.1 instead of the 36.9 they would buy on white noise.

A sign, and what it cost

The sequence route needed a fix before it could be used, and the fault is worth recording because it is invisible to every check except the one that was eventually run.

Marching a sequence through the pink netlist with the nodal solver takes twelve seconds for sixty thousand samples, which is the wrong price for a figure behind a slider. So each lag section is converted to its bilinear equivalent and the sequence is filtered directly — same poles, same zeros, same design, two arithmetics. Collecting the transform gives

y[n](1+kp)+y[n1](1kp)=x[n](1+kz)+x[n1](1kz)y[n](1+k_p) + y[n-1](1-k_p) = x[n](1+k_z) + x[n-1](1-k_z)

so the past output moves to the right-hand side with its sign reversed. It shipped without the reversal. For a low-frequency pole kₚ is large, so the coefficient that should have been close to +1 was close to −1, and a pole at z ≈ −1 is a filter that alternates every sample instead of integrating over many — the exact opposite of what a 1/f shape is.

What made it findable is that the result was not merely wrong but impossible. The sequence came out with a sample-to-sample standard deviation of 10⁴ and a four-sample block mean of 2.2: a ratio of four thousand from a four-sample average, which no averaging of real data can produce. The averaging exponent read −0.89, and −0.89 is steeper than white — a claim that averaging works better on correlated noise than on independent noise, which nothing in the subject would support.

This is the second number in this phase that was recorded as verified and did not reproduce. Neither had an assertion holding it. The exponents in every figure here are now asserted at every position of the slider, and one of the assertions is precisely that the two routes differ by many times the seed-to-seed spread — which the broken filter would have passed, and which the check below it, that white comes out at −0.5, would not.

What the measurement says

Both sequences come from one seeded white stream: the pink one is that stream put through the network above, so the comparison is of a sequence against itself filtered, and no second random draw enters anywhere. Averaged in non-overlapping blocks of 1, 4, 16, 64, 256 and 1024 samples, across five seeds:

exponent what 1024 samples buy
white n0.510±0.006n^{-0.510 \pm 0.006} a factor of 36.9
pink n0.087±0.013n^{-0.087 \pm 0.013} a factor of 2.1

The white result is the √N law measured rather than quoted, and it is the control: a route that could not reproduce −0.5 on white noise would have nothing to say about pink. The pink result is the essay. Averaging a flicker sequence for a thousand times as long improves it by a factor of two.

The blocks are non-overlapping deliberately. Overlapping blocks share samples, so the estimates being compared would be correlated with each other on top of the samples inside them being correlated, and the exponent would be measuring the overlap as much as the noise.

The first fifth of every run is discarded, and that is not tidiness either. The slowest section’s pole is at the bottom of the band and its step response is still arriving many thousands of samples in. Measured over the transient, the first block’s standard deviation comes out four orders of magnitude above the rest — a settling curve reported as a spread.

Why it happens, in one sentence about time

The √N law is a statement about independent samples, and its proof is the addition of variances: N independent things of variance σ² sum to variance ², so their mean has variance σ²/N and their spread falls as 1/√N.

Correlated samples do not add that way. If two samples are correlated, the variance of their sum includes a cross term that does not vanish, and the average is worse than the independent count suggests. The size of the effect is set by how long the correlation lasts relative to the block: a noise whose correlation dies away in a microsecond is effectively independent over a millisecond block, and the law holds.

Flicker noise has no such time. A spectrum going as 1/f down to the lowest frequency measured means fluctuations on every time scale, in equal power per decade — as much wander over a day as over a second. There is no averaging time long enough to be long compared with the correlation, because the correlation extends to whatever time is chosen.

That is what the slider in the first figure demonstrates directly. Raising the bottom of the 1/f band from 0.5 Hz to 40 Hz gives the noise a longest time scale, and the exponent recovers from −0.083 to −0.172 while the thousand-sample benefit goes from 2.1× to 3.7×. It never gets back to −0.5, because the band is still 1/f over the block lengths being used; but the direction is unambiguous and the mechanism is visible.

Averaging a white sequence, and averaging a pink one. computed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.172 ± 0.004, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 3.7 on the second.
Fig. 4 The same measurement with the 1/f region starting at 40 Hz instead of 1 Hz. The blocks now reach below the bottom of the band, the sequence has a longest correlation time again, and averaging partially recovers — n0.172n^{-0.172} and a factor of 3.7 for a thousand samples. Still not the √N law, and no longer nothing.

Equal power per decade, which is the whole shape in four words

There is a compact way to say what 1/f means, and it explains both the difficulty and why the construction above works.

Integrate a flat density over a band and the power is proportional to the width of the band, so a decade from 1 kHz to 10 kHz carries a hundred times the power of a decade from 10 Hz to 100 Hz. That is what “white” means, and it is why the noise field’s earlier essays could work in bandwidths at all: almost all the noise is at the top.

Integrate 1/f over a band and the power is proportional to the logarithm of the ratio of its ends, so every decade carries the same power. 0.1 Hz to 1 Hz carries exactly as much as 1 kHz to 10 kHz. A measurement that runs for ten times as long does not average an existing amount of noise down; it opens a new decade at the bottom and admits a fresh helping of it, and the two effects very nearly cancel. That near-cancellation is the −0.087 exponent, arrived at from the spectrum instead of from the sequence.

It is also why the integral of a pure 1/f density has no value until both ends are named. Neither end can be dropped: at the top the flicker region gives way to the white floor, and at the bottom nothing has ever been observed to give way to anything. Every flicker figure quoted anywhere carries a band, whether or not it is printed beside the number.

What the standard deviation stops being

The awkwardness runs deeper than slow convergence, and it has a name in the fields that had to deal with it first.

For an independent sequence, the sample standard deviation is an estimate of a quantity that exists: take more data and the estimate settles down. For a 1/f sequence the underlying quantity is not there to be estimated — the variance of the process depends on how long it was observed, because a longer observation opens lower decades — so a longer run does not refine the answer, it changes the question.

The block table above shows this happening. Each row is a spread computed over blocks of a stated length, and for the white sequence those spreads are estimates of one number scaled by √n. For the pink sequence they are not estimates of anything in common: they are six different measurements of six different processes, distinguished by the time window each of them integrated over.

This is why frequency metrology abandoned the standard deviation altogether and uses the Allan variance instead — the mean squared difference between successive averages rather than the spread about a global mean. Differencing adjacent blocks cancels the low-frequency wander that the global mean cannot, and the result converges for noise processes where the ordinary variance does not. The same construction, under other names, is what a lock-in amplifier, a chopper and a correlated double sampler all do.

Where the shape comes from, and why the essay does not derive it

This essay builds a 1/f density and measures its consequences; it does not explain it, and the reason is worth being direct about.

There is no single accepted mechanism. The most durable account is that a superposition of trapping-and-release processes with exponentially distributed activation energies produces a superposition of Lorentzians whose corner frequencies are spread uniformly in the logarithm — and a uniform-in-log spread of corners integrates to 1/f, which is precisely what the interleaved chain in the figure does deliberately. The chain is a model of that account, at four or sixteen sections rather than at Avogadro’s number of them, and the visible staircase is the difference.

What makes it unsatisfying as an explanation is that the same spectrum appears in things that share no mechanism at all. The site’s rule is that a model is drawn with the condition under which it stops being true, and the honest statement here is that the 1/f shape is an observation with a construction that reproduces it, not a derivation. Everything in this essay is a consequence of the shape, and every consequence survives whatever the shape turns out to come from.

Averaging a white sequence, and averaging a pink one. computed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.108 ± 0.007, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 2.3 on the second.
Fig. 5 Five hertz: exponent −0.108, and 1,024 samples buy 2.3 times. Moving the corner up by a factor of ten has improved the averaging by ten per cent, which is the useful form of the bad news — the benefit of averaging is set by how far the record extends below the corner, and a record is short.
Averaging a white sequence, and averaging a pink one. computed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.126 ± 0.006, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 2.6 on the second.
Fig. 6 And ten hertz: −0.126, 2.6 times. Across the four corners drawn the pink exponent runs −0.083, −0.108, −0.126 and further at 40 Hz, against the white sequence’s −0.510 at every one of them. Averaging is not a technique that works less well on pink noise; it is a technique that has almost stopped working, and the number that says so is an exponent rather than a factor.

What is done about it, and what is not

Two responses exist and the difference between them is the point.

Chopping is the response that works. The signal is modulated up to a frequency above the flicker corner before the amplifier sees it, amplified there, and demodulated back down afterwards. The amplifier’s own 1/f noise is at baseband and is not modulated, so it lands away from the signal and is filtered out. The signal has spent its time in the amplifier at a frequency where the density is flat. This is not averaging harder; it is moving the measurement to where averaging still works.

Averaging longer is the response that does not. It is the intuitive one, it is what an instrument does when it is asked for more resolution, and the figures above are what it buys: a factor of two for a factor of a thousand in time. Beyond the flicker corner, patience is very nearly worthless, and the second decimal place of a slow measurement is bought with a technique rather than with a clock.

There is a third thing, which is neither, and it is the one this field has already drawn. The resistor’s own 4kTR is white all the way down — thermal noise has no flicker region, because it is a property of temperature and resistance and neither has a memory. So the floor a resistor sets can be averaged down indefinitely, and the floor a circuit has cannot. The two essays of this ladder are separated by exactly that.

There is a fourth response that is neither technique nor patience, and it is the one the previous essay’s arithmetic already covers: choose the part. The flicker corner is a device parameter like eₙ and iₙ, it varies by two orders of magnitude between technologies, and a field-effect input with a corner at 10 Hz and a bipolar input with a corner at 1 kHz are different propositions for a slow measurement even when their flat-band densities are identical. The datasheet’s front-page number describes the flat region; the corner describes how much of the measurement is in it.

Which is the sentence this ladder’s two rungs come down to. The first rung said an amplifier’s noise is a property of the device and of the source it is given. The second says that both halves of that statement have a frequency attached, and that below the corner the useful ways of improving a measurement stop being the obvious ones.

The rule this model stops being true under

Every figure on this site carries the frequency, amplitude or size at which its model stops applying, and for the √N law the quantity is a time.

The honest version is that there is no such time, and that is the finding. What the measurement gives instead is a rate of decay: n0.087n^{-0.087} against n0.510n^{-0.510}, on one stream, filtered and not. The familiar law is not merely degraded here; it is reduced to a factor of two for three orders of magnitude of patience, and the useful conclusion is that a measurement limited by flicker noise has to be redesigned rather than repeated.

What redesigning means, and where the collection does it

“Redesigned rather than repeated” is a conclusion that needs a technique attached to it, and this collection has the technique in a field that arrived at it for another reason.

The sample that is subtracted is correlated double sampling: the reset level a capacitor holds is the same number in two consecutive samples and cancels exactly. What makes that relevant here is why it cancels — the two samples are close together in time, so any disturbance slower than the interval between them is common to both and subtracts. A flicker process is precisely a disturbance whose power is at low frequencies, so subtracting two nearby samples is the one operation that removes it, and it does so for the same reason that averaging does not.

The price is the one that essay measures: the amplifier’s white noise does not cancel, since two samples of it are independent, so its variance doubles. Which is exactly the trade this page implies — a technique that helps against the correlated part of the noise and hurts against the uncorrelated part — and it says where the crossing is: at 3 dB of penalty on the white component against however many decibels the flicker component was costing.

The bandwidth noise sees supplies the other half of the practical arithmetic, since both the averaging law and its failure are statements about a bandwidth: a noise computed from a −3 dB point is twenty-one per cent low for a single pole and four per cent high for a five-pole Chebyshev. A measurement being redesigned around its own flicker corner is a measurement whose effective bandwidth is being chosen deliberately, and that essay is what turns the choice into a number.

Two other measurements in this collection rest on samples being independent, and both say so. Two thresholds because there is a floor counts transitions over a record and finds the count proportional to the record length — which is what an independent-samples argument predicts and which a flicker component would break, since a slow wander carries the signal back across the threshold in fewer, larger excursions. And the floor a resistor sets is the floor this essay’s white sequence is drawn from: 4.00 nanovolts per root hertz, thermal, with no memory in it at all.

Which is the practical division. Everything with 4kTR4kTR in it averages as N\sqrt N; everything with a device’s own flicker corner in it does not, and the corner is a property of the part rather than of the circuit around it.

What the corner does to the field’s other measurements

Averaging is how nearly every number in this field is measured, so a corner below which it stops working is a statement about the field’s own instruments. The floor a resistor sets is the two-route measurement that the pink case would defeat if the record were long enough. The floor a circuit has is where the amplifier’s own 1/f corner sits relative to a measurement bandwidth. A floor and a ceiling is the dynamic range whose lower end this decides, and The bandwidth noise sees is the factor that assumes a white density. One step, computed twice is the discipline the marched half of this page belongs to, and The total that has no resistor in it is the one floor in the field that no amount of averaging would improve anyway.

Part 2 on device noise

One argument about Device noise, and one of 3 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 20.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

AveragingBilinear transformConvergence orderCorrelation timeNoise figureFlicker noise