The floor, which bounds from below

The sample that is subtracted

Three rungs of this argument have measured floors that no gain moves and no filter reaches, because both arrive as numbers already sampled. One of them can be subtracted: the reset level a capacitor holds is the same number in two consecutive samples and cancels exactly. What it costs is that the amplifier's own noise is not — two samples of it are independent, so its variance doubles. That is thirty times better at a megahertz, a loss above sixty, and the crossing is the one the rung below computed for a different question.

Assumes: The total that has no resistor in it · The floor a resistor sets · The bandwidth noise sees

The three rungs below this one are about floors that cannot be filtered.

The first measured kT/CkT/C and found it contains no resistance, no clock and no capacitor ratio: switch a capacitor and what is left on it has a mean square of kT/CkT/C, and nothing about the switch appears anywhere in it. The second found that making the clock slower or faster does not change it. The third put an amplifier in the same loop and found that the number of times its noise folds into the band is exactly the number of time constants the settling needs — so the amplifier’s term rises with the clock while kT/CkT/C does not, and the two cross at 60.2 MHz for a picofarad.

All three are unfilterable for the same reason: they arrive as a number that has already been sampled, and a filter downstream of a sampler cannot separate what folded on top of what.

One of the two can be subtracted.

Subtracting removes kT/C entirely and doubles the amplifier — worth 31× at a megahertz and a loss above 60 MHz. computed by solving, not by drawing. The noise on one sample of a switched-capacitor stage, and on the difference of two samples taken a settled interval apart, against clock frequency. The reset level is the same number in both samples and cancels exactly; the amplifier's own noise is two independent samples and its variance doubles, measured at 2.000 against the 2 the correlation predicts. At a megahertz that is 63.8 µV down to 11.53 — 31 times in power. The two curves cross at 60.2 MHz, which is where the amplifier's own noise equals kT/C, and above it the subtraction costs more than it removes.
Fig. 1 The noise on one sample of a switched-capacitor stage, and on the difference of two samples taken a settled interval apart, against clock frequency. One floor is removed exactly; the other is doubled.

Why one cancels and the other does not

The reset level is a number. When the switch opens, the capacitor is left holding a particular voltage — a draw from a distribution whose mean square is kT/CkT/C — and from that instant until the switch closes again, that number does not change. It is not a process, it is a value.

So sample the capacitor before the signal arrives, sample it after, and subtract. The reset level appears identically in both samples and vanishes from the difference. Not reduced, not filtered, not averaged down: the same number, subtracted from itself.

The amplifier’s own noise is not a number. It is a continuous process filtered by the settling response and sampled, and two samples of it taken at different instants are two different draws. Subtracting them does what subtracting two independent random variables does — the variances add:

var(v2v1)=2σ2(1ρ),ρ=eΔt/τ\mathrm{var}(v_2 - v_1) = 2\sigma^2(1 - \rho), \qquad \rho = e^{-\Delta t/\tau}

For an interval much longer than the settling time constant, ρ\rho is nothing and the penalty is exactly two in power, 2\sqrt2 in root-mean-square.

Five resistances, five corner frequencies, and one total on the capacitor. computed by solving, not by drawing. The noise density at a 1 pF capacitor charged through 100 Ω, 1 kΩ, 10 kΩ, 100 kΩ and 1 MΩ, at 290 K. The densities are 100 times apart and the noise bandwidths 1.0e+4 times apart, in opposite directions, so the area under every curve is the same: 63.2762 µV against 63.2762 µV, and √(kT/C) is 63.2762 µV. The resistance has cancelled out of the answer, and the reason is that ½C⟨v²⟩ is the ½kT a degree of freedom in contact with a bath holds — which no arrangement of resistors can change. The claim is about the whole frequency axis and nothing less: inside a 15.9 MHz band the same five networks give 5.05 µV to 63.07 µV, a factor of 12.5.
Fig. 2 The floor being removed, from the first rung of this argument: a mean square of kT/C with no resistance in it.

What that is worth, and where it stops being worth it

For a picofarad, a 4 nV/√Hz part and twelve bits of settling at a one-megahertz clock:

  • one sample carries 63.8 µV rms, of which almost all is kT/CkT/C;
  • the difference of two carries 11.53 µV;
  • which is 30.6 times in power.

Two independent routes give that. The closed form above, and a seeded white sequence marched through the settling exponential as a difference equation, given a fresh reset offset once per cycle, sampled twice and subtracted. They agree to 0.035 per cent on a measurement whose own spread is 1.0 per cent, and they share nothing: one is an integral of a correlation function and the other is forty thousand pseudo-random numbers.

The march is where the cancellation is demonstrated rather than argued. The reset offset is a real number added to both samples, and it disappears from the difference because the arithmetic makes it disappear, not because anybody told the simulation to remove it.

Every bit of settling costs 0.693 of a fold, and the noise the square root of it. computed by solving, not by drawing. The time constants a switched-capacitor stage must settle through in half a clock period, against the accuracy asked for, with the fold count drawn on the same axis because they are the same number. Each extra bit is ln 2 = 0.6931 more time constants exactly, so the amplifier's sampled noise rises as the square root of the accuracy demanded: 6.66 µV at eight bits and 9.42 µV at sixteen on a 1 pF capacitor clocked at 1.00 MHz, a factor of 1.41. √(kT/C) is 63.3 µV and knows nothing about any of it.
Fig. 3 The rung below, which is the other half of this trade: the number of times a sampled amplifier’s noise folds into the band is exactly the number of settling time constants the bit count needs.

Now the part that makes it a trade. The subtraction removes a floor that does not depend on the clock and doubles one that rises with it. So above some clock frequency, twice the amplifier’s noise exceeds the amplifier’s plus kT/CkT/C, and subtracting is a loss.

That crossing is where σamp2=kT/C\sigma_\mathrm{amp}^2 = kT/C — which is exactly the crossing the rung below computed when it asked which of the two floors dominates. 60.2 MHz for these numbers, and the swept measurement changes hands between 46 and 68 MHz. One frequency answers two different questions about the same stage, and neither question mentions the other.

Above it the picture inverts completely: at a gigahertz clock, one sample carries 266 µV and the difference carries 365. The arrangement that removes the larger floor at low speed is adding to the larger floor at high speed.

Subtracting removes kT/C entirely and doubles the amplifier — worth 4× at a megahertz and a loss above 6 MHz. computed by solving, not by drawing. The noise on one sample of a switched-capacitor stage, and on the difference of two samples taken a settled interval apart, against clock frequency. The reset level is the same number in both samples and cancels exactly; the amplifier's own noise is two independent samples and its variance doubles, measured at 2.000 against the 2 the correlation predicts. At a megahertz that is 21.6 µV down to 11.53 — 4 times in power. The two curves cross at 6.0 MHz, which is where the amplifier's own noise equals kT/C, and above it the subtraction costs more than it removes.
Fig. 4 Ten picofarads, where kT/C is ten times smaller in power and the crossing falls to 6.0 MHz. A larger capacitor makes the reset level less worth removing, which is the same statement as the rung below’s “above the crossover a larger capacitor buys nothing at all”, read from the other side.

Why this is not filtering, and why that matters

It is worth being precise about what kind of operation this is, because “subtract two samples” and “filter the noise” sound like the same sentence and the difference is the whole reason the arrangement works.

A filter placed before the sampler can remove noise above half the clock, and the rung below measured exactly what that is worth: the amplifier’s noise folds as many times as its settling needs time constants, and the folding happens because the settling bandwidth has to be wide enough for the signal. A filter narrow enough to stop the folding is a filter the signal cannot get through.

A filter placed after the sampler cannot separate anything, because the folding has already happened and every frequency that folded is now indistinguishable from every other.

The subtraction is neither. It is an operation on two samples of the same held quantity, and what it exploits is not a frequency separation at all but the fact that one of the two contributions is a constant over the interval and the other is not. That is a statement about time, not about spectrum, and it is the only handle available on a floor that is already sampled.

The capacitor disappears

Sweeping the sampling capacitor at a fixed clock puts the whole of the argument in one table:

capacitor one sample subtracted in power
0.1 pF 200.3 µV 11.53 µV 301×
0.3 115.8 11.53 101×
1 63.80 11.53 30.6×
3 37.43 11.53 10.5×
10 21.61 11.53 3.5×

The right-hand column is the same number five times.

That is not a rounding: after the subtraction the capacitor is not in the answer at all. What is left is the amplifier’s own noise doubled, and the amplifier’s own noise is en2/4τe_n^2/4\tau with τ\tau set by the clock and the bit count. The capacitance appears nowhere in it.

The first rung of this argument is titled for a total that has no resistor in it. The fourth is about a floor that has no capacitor in it, and the two facts have the same cause read from opposite ends: the reset level’s mean square is kT/CkT/C precisely because the noise bandwidth of the switch’s own resistance is 1/4RC1/4RC and the RR cancels — and once that term is subtracted away, the only term left is one the capacitor never entered.

The design consequence is the sharpest thing in this essay. Without the subtraction, a stage’s floor is improved by making the capacitor larger, which costs area, costs settling current in proportion, and is the reason a low-noise switched-capacitor stage is a large one. With the subtraction, a tenth of a picofarad is exactly as quiet as ten, and the whole of that expenditure buys nothing.

Subtracting removes kT/C entirely and doubles the amplifier — worth 301× at a megahertz and a loss above 602 MHz. computed by solving, not by drawing. The noise on one sample of a switched-capacitor stage, and on the difference of two samples taken a settled interval apart, against clock frequency. The reset level is the same number in both samples and cancels exactly; the amplifier's own noise is two independent samples and its variance doubles, measured at 2.000 against the 2 the correlation predicts. At a megahertz that is 200.3 µV down to 11.53 — 301 times in power. The two curves cross at 601.7 MHz, which is where the amplifier's own noise equals kT/C, and above it the subtraction costs more than it removes.
Fig. 5 A tenth of a picofarad, where one sample carries 200 µV and the difference carries the same 11.53 µV as every other capacitance. The crossing moves out to 602 MHz, because kT/C is ten times larger in power and takes longer for the amplifier to catch.

The interval is the settling time, and the two demands pull opposite ways

The doubling is not a constant. It is 2(1eΔt/τ)2(1 - e^{-\Delta t/\tau}), and Δt\Delta t is a design quantity: the interval between the reset sample and the signal sample.

Take the two samples close together and the amplifier’s noise is correlated across the interval, so part of it cancels along with the reset level. Measured, at a fifth of a settling time constant the penalty is 0.307 rather than 2 — the subtraction is removing five sixths of the amplifier’s noise as well.

That looks like a free improvement until the other use of the same interval is remembered. The interval between the two samples is the interval the signal has to settle in. A stage settling to twelve bits needs ln212=8.32\ln 2^{12} = 8.32 time constants, and at 8.32 time constants the correlation is 2×1042\times10^{-4} and the penalty is 2.000 to four figures.

The doubling is 2(1 − e^(−Δt/τ)), so the interval that saves noise is the one the signal needscomputed by solving, not by drawing. The variance of the difference of two samples, against one sample, as a function of the interval between them in settling time constants. At 33 time constants the two samples are independent and the penalty is exactly two; at 0.17 it is 0.307, because the amplifier's noise is correlated across that interval and part of it cancels along with the reset level. The marched sequence is drawn over the closed form and agrees to 3.8 per cent. The trade is that this interval is the one the signal has to settle in: 12 bits needs 8.32 time constants, which is past where the correlation has gone.012interval between the two samples, in settling time constantsamplifier noise power, against one sample0.170.330.832368173312 bits of settlingtwo independent samplesbits12settling needs8.32 τat 0.1 τ0.307×at 1 τ1.264×settled2.000×worst two-route gap3.84%solved, then checked — a correlation, marched and in closed form0.31× at 0.17 τ, 2.00× settled
Fig. 6 The penalty against the interval, in settling time constants, with the marched sequence drawn over the closed form. The vertical line is where twelve bits of settling puts the interval, and it is past everywhere the correlation was doing anything. Drag it through the bit count.

So the two demands are on one number and they are not close: the region where the correlation helps is under about two time constants, and any stage that settles to more than three bits is past it. The penalty is 2 for every converter anybody builds, and the 2(1eΔt/τ)2(1 - e^{-\Delta t/\tau}) is worth knowing because it says why it is 2 rather than because it is ever anything else.

There is one arrangement in which it is not, and it is worth naming: a stage that samples the reset level and the signal at the two ends of a short interval within a long clock period — which is what a correlated double sampler in an image sensor does, where the two levels are a microsecond apart on a pixel read out over tens. There the interval and the settling are genuinely different quantities and the correlation is available.

The doubling is 2(1 − e^(−Δt/τ)), so the interval that saves noise is the one the signal needs. computed by solving, not by drawing. The variance of the difference of two samples, against one sample, as a function of the interval between them in settling time constants. At 44 time constants the two samples are independent and the penalty is exactly two; at 0.22 it is 0.398, because the amplifier's noise is correlated across that interval and part of it cancels along with the reset level. The marched sequence is drawn over the closed form and agrees to 4.4 per cent. The trade is that this interval is the one the signal has to settle in: 16 bits needs 11.09 time constants, which is past where the correlation has gone.
Fig. 7 Sixteen bits, where the settling needs 11.09 time constants. The line moves right and the useful region does not move at all, so the deeper the converter the more completely the penalty is exactly two.

What the arrangement does to a flicker floor

One more thing the subtraction does, and it is the reason image sensors use it rather than converters.

The difference of two samples Δt\Delta t apart is a filter, and its magnitude response is 1ej2πfΔt=2sin(πfΔt)|1 - e^{-j2\pi f\Delta t}| = 2|\sin(\pi f \Delta t)|. At low frequency that is 2πfΔt2\pi f\Delta t — it falls to nothing as the frequency does.

So the subtraction is a high-pass, and everything the amplifier does slowly is removed along with the reset level: its offset, its drift, and the low-frequency part of its own noise. For a part whose noise rises as 1/f1/f below some corner, the contribution from below the corner is suppressed by the square of the frequency, which is exactly the shape that makes flicker noise unfilterable by ordinary means.

This collection has measured the corner where averaging stops working, and the reason it stops is that 1/f1/f noise has as much power in every decade — so waiting longer does not help. A subtraction of two samples does not wait longer; it discards everything slower than 1/Δt1/\Delta t by construction.

Averaging a white sequence, and averaging a pink one. computed by solving, not by drawing. Both sequences are the same seeded white stream, one of them put through the 1/f network. Averaged in non-overlapping blocks, the white one's spread falls as n to the -0.510 ± 0.006 across five seeds — the √N law — and the pink one's as n to the -0.087 ± 0.013, which is very nearly not at all. A thousand-sample average buys a factor of 36.9 on the first and 2.1 on the second.
Fig. 8 The floor this arrangement removes and averaging cannot, from the noise field’s own essay: a white sequence’s spread falls as the square root of the block length and a 1/f sequence’s barely falls at all.

What the difference of two samples really is

The closed form used above is a correlation of an exponentially filtered process, and it is worth one paragraph in the other language, because the two say the same thing and the second is the one that generalises.

Subtracting two samples Δt\Delta t apart is a filter whose response to a frequency ff is 1ej2πfΔt1 - e^{-j2\pi f\Delta t}, of magnitude 2sin(πfΔt)2|\sin(\pi f \Delta t)|. The output variance is therefore

0S(f)H(f)24sin2(πfΔt)df\int_0^\infty S(f)\,|H(f)|^2\,4\sin^2(\pi f\Delta t)\,\mathrm{d}f

with SS the source’s density and HH the settling response. The 4sin24\sin^2 term averages to 2 across any band wide compared with 1/Δt1/\Delta t, which is the factor of two again; it is (2πfΔt)2\approx (2\pi f\Delta t)^2 at the bottom, which is the high-pass; and it is 4 at f=1/2Δtf = 1/2\Delta t and its odd multiples, which is the part nobody mentions.

That last one is a real feature and not a curiosity: a subtraction of two samples amplifies the noise at half the reciprocal of their spacing, by six decibels. For a source with a peak there — a switching supply’s ripple, a clock harmonic, another converter on the same die — the arrangement makes it worse, and the frequency it makes worse is set by a timing choice rather than by anything electrical.

What a real double sampler adds

The arrangement measured here has two samples and one capacitor. A built one has two capacitors, because the first sample has to be stored somewhere while the second is taken, and the store is not free.

The second capacitor is reset too, so it contributes a kT/CkT/C of its own that the subtraction does not remove — it appears in one sample and not the other. Whether that matters depends on whether it is reset once per cycle or held across many, which is an arrangement question rather than a noise one, and the honest statement is that the floor measured in this essay is the floor of the idea rather than of any built circuit.

Two other things a real one has. The switch that opens to take the first sample injects charge onto the node, and that injection is removed by the subtraction only if it is identical at both samples — which it is not, because it depends on the voltage the switch is at. And the two samples pass through the same amplifier at two different times, so anything that has changed in between, including the amplifier’s own recovery from the first sample, is signal.

Where the floor this removes came from

Correlated double sampling removes one floor and leaves several. The noise a clock does not make is the kT/C floor it does not remove, which belongs to the capacitor and not to the switch. A resistor made of a clock is the circuit both floors live in. The floor a resistor sets is where the two-route discipline behind every number here is established, and The bandwidth noise sees is the factor that turns a density into a voltage. The amplifier inside the sample is the floor that takes over once the clock is fast enough, and this page’s subtraction does nothing about it at all.

What is checked

The two routes are required to agree on the subtracted variance to eight times the estimate’s own spread, and the spread is computed from the record length rather than chosen — which is what makes the agreement a check rather than a coincidence of tolerances.

The penalty at a settled interval is asserted to be exactly two to a part in a thousand, and at a short interval to be under a half — the two ends that establish the correlation is real. That the marched and closed-form penalties agree at every interval is asserted too, including the short ones, which needed the march to be run faster: at sixty-four steps per clock an interval of a hundredth of a clock is one step, and a first version reported a fifty per cent error that belonged entirely to the rounding of the interval rather than to anything about noise.

That the reset level is gone rather than reduced is asserted directly: the subtracted variance is required to be under twice the amplifier’s own plus a millionth of kT/CkT/C, so a version that filtered the reset level rather than cancelling it would fail.

And the trade is asserted as a trade — some clocks on the sweep must improve and some must get worse, and the frequency at which it changes hands is asserted against the rung below’s crossover rather than against a number.

What is not modelled: the switch’s own thermal noise during the acquisition, which is what sets kT/CkT/C in the first place and is already in it; the charge injected by the switch as it opens, which is a signal-dependent offset that this arrangement removes only if it is the same at both samples; the second capacitor a real double sampler needs, which contributes a reset level of its own; and the flicker suppression, which is argued here from the shape of the filter and is measured nowhere in this essay.

Why the crossing is the same crossing

The frequency at which this technique stops paying is the one the amplifier inside the sample computed for a different question, and that is not a coincidence worth passing over.

That essay asks which of two floors dominates — the switches’ kT/CkT/C, which no design choice moves, against the amplifier’s white noise folded as many times as the settling has time constants — and finds the switches below sixty megahertz and the amplifier above. This essay asks which of two arrangements wins, and the answer is decided by the same comparison read backwards: correlated double sampling removes exactly the floor that dominates below the crossing and doubles exactly the one that dominates above it. So the technique is thirty times better at a megahertz and a loss above sixty, and the sixty is the same number because it is the same pair of quantities being compared.

That is a satisfying kind of result and it is also a warning about how the number transfers. The crossing is a property of one capacitor, one amplifier and one settling requirement, so a design with a larger sampling capacitance has a lower kT/CkT/C and a lower crossing — meaning the arrangement that is obviously worth having on one part is obviously not on another built for a different resolution. The total that has no resistor in it is where the first of those two quantities is established, and it contains nothing but kk, TT and the capacitor, which is what makes the crossing computable before anything is built.

There is a second reason to reach for this arrangement that has nothing to do with kT/CkT/C, and it is the one that makes the technique worth having even above the crossing. The corner where averaging stops working measures what happens to a flicker process under averaging — a block mean falling as the 0.087-0.087 power of the block length against white noise’s 0.510-0.510, so a thousand samples buy a factor of 2.1 rather than 36.9 — and a subtraction of two nearby samples is the one operation that does work against it, because anything slower than the interval between them is common to both. So the same doubling of the white variance that makes this technique a loss above sixty megahertz is bought with a removal of an amplifier’s flicker noise that no amount of averaging would have achieved.

Part 4 on kt over c

One argument about Kt over c, and one of 4 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Correlation timeDesign tradeoffJohnson noiseKt over cModel rangeNoise bandwidthSettling timeVerification