Where a signal becomes a number

One knob, and the two exponents it turns

Oversampling is quoted as buying one thing and buys two that improve at different rates. The hold's droop at the band edge falls with the SQUARE of the ratio — fitted exponent −2.011 over six doublings, from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it — while the nearest image moves out with the first power, exponent 1.022. The third quantity, the poles a reconstruction filter needs for sixty decibels, collapses from 20.5 to 3.21 by a ratio of four alone, because it is a logarithm of the second.

Assumes: The staircase on the way out · The frequency a sample rate invents

Three earlier essays here have measured consequences of one rectangle and every one of them has ended at the same place: the answer is a function of rr, the band edge as a fraction of the clock, and rr is the one quantity a designer controls.

The staircase on the way out found a droop of 20logsinc(πr)20\log\mathrm{sinc}(\pi r). The nulls are where nothing is found an image rejection of sinc(πr)/sinc(π(1r))\mathrm{sinc}(\pi r)/\mathrm{sinc}(\pi(1-r)) and a filter transition ratio of (1r)/r(1-r)/r. Flatness, and the two currencies it is bought in found two prices for correcting the first, both functions of the same rr.

Lowering rr is oversampling, and it is normally described as buying a gentler filter. It buys at least three separate things, and they improve at three different rates. The rates are what decide which of them a given ratio is actually being bought for, and none of the three is usually measured against the others on one axis.

What an oversampling ratio buys, and at two different ratescomputed by solving, not by drawing. A 20 kHz band on a 48 kHz base clock, interpolated by ratios from 1 to 64. The hold's droop at the band edge falls with the SQUARE of the ratio — the fitted exponent over six doublings is -2.0113 — from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it. The nearest image moves out with the FIRST power, exponent 1.0222, from 1.40 times the band edge to 37.4. So one decision buys two things at rates differing by a factor of two in the exponent, and the third quantity — the poles a reconstruction filter needs for sixty decibels — collapses from 20.5 to 3.21 by a ratio of four alone.1m10m100m110100110oversampling ratiodecibels of droop; the image's distance as a multiple of the band edgedroop (dB), slope -2.011image distance, slope 1.022poles for 60 dBdrawn hereratio×4clock192 kHzdroop at 20 kHz0.1556 dBnearest image172 kHzwhich is8.60× the bandpoles for 60 dB3.21droop exponent-2.0113image exponent1.0222solved, then checked — two exponents from one knobdroop -2.01, images 1.02
Fig. 1 A 20 kHz band on a 48 kHz base clock, interpolated by ratios from one to sixty-four. Three quantities on one logarithmic pair of axes: the droop at the band edge in decibels, the nearest image’s distance as a multiple of the band edge, and the number of maximally flat poles needed for sixty decibels between them. The slider is the ratio.

Three quantities, and their exponents fitted over six doublings

The arrangement is the ordinary one. A signal band that stops at 20 kHz, a base clock of 48 kHz, and an interpolator that raises the sample rate by LL before the converter. Raising the rate does not change the band and does not change the samples’ information; it changes where everything sits relative to the clock.

ratio clock droop at 20 kHz nearest image poles for 60 dB
×1 48 kHz 2.640 dB 1.40× the band 20.53
×2 96 kHz 0.6292 dB 3.80× 5.17
×4 192 kHz 0.1556 dB 8.60× 3.21
×8 384 kHz 0.0388 dB 18.20× 2.38
×16 768 kHz 0.0097 dB 37.40× 1.91
×32 1.536 MHz 0.0024 dB 75.80× 1.60
×64 3.072 MHz 0.0006 dB 152.6× 1.37

Read down the droop column: each doubling divides it by four. Read down the image column: each doubling roughly doubles it. Read down the third: each doubling takes about a fifth off it, and the fifth gets smaller.

Fitted over the whole range rather than read off two points — a fit over two points is an arithmetic identity and proves nothing — the exponents come out at −2.011 for the droop and 1.022 for the image separation. Both are checked against their integer values, and the second is checked over the upper half of the range for a reason given below.

The origin of the first is a series. Near the origin the sinc is 1(πr)2/61 - (\pi r)^2/6, so the decibels are 20ln10(πr)26\tfrac{20}{\ln 10}\cdot\tfrac{(\pi r)^2}{6} to leading order, and rr is inversely proportional to LL. The droop is therefore quadratic in the ratio and there is nothing to choose about it.

The origin of the second is subtraction. The nearest image sits at LfsfbandLf_s - f_{\text{band}}, so its distance from the band edge is Lfs/fband1Lf_s/f_{\text{band}} - 1 — proportional to LL with an offset. The offset is not negligible at small LL: at the Nyquist rate the image is at 1.40 times the band edge rather than at 2.40, because the subtraction has taken away most of it.

What an oversampling ratio buys, and at two different rates. computed by solving, not by drawing. A 20 kHz band on a 48 kHz base clock, interpolated by ratios from 1 to 64. The hold's droop at the band edge falls with the SQUARE of the ratio — the fitted exponent over six doublings is -2.0113 — from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it. The nearest image moves out with the FIRST power, exponent 1.0222, from 1.40 times the band edge to 37.4. So one decision buys two things at rates differing by a factor of two in the exponent, and the third quantity — the poles a reconstruction filter needs for sixty decibels — collapses from 20.5 to 3.21 by a ratio of four alone.
Fig. 2 No interpolation at all, which is the case the whole argument is usually stated for and is the one nobody builds. The droop at 20 kHz is already 2.64 dB on a 48 kHz clock, the nearest image is 1.40 times the band edge away, and a maximally flat filter needs twenty and a half poles to put sixty decibels between them.

The exponent that is asymptotic, and where it says so

The first power is a limit rather than a law, and the rule here is to say where a law stops being one.

Over the full range from one to sixty-four the fitted exponent for the image distance is 1.109, not 1.0. Over the upper half — eight to sixty-four — it is 1.022. The difference is eleven per cent and it is entirely the 1-1 in Lfs/fband1Lf_s/f_{\text{band}} - 1, which matters when LL is small and does not when it is large.

Eleven per cent would be a rounding error if the small-LL end were an academic corner. It is not: a converter with no interpolator at all sits at L=1L = 1, and a great many instrument outputs, arbitrary waveform generators and control DACs do exactly that. So the region where the first power is wrong by eleven per cent is precisely the region a designer without an interpolating filter is working in, and the correct statement there is the subtraction rather than the power law.

The droop’s exponent has the same structure and is better behaved. Its series is in (πr)2(\pi r)^2 and the next term is (πr)4/120(\pi r)^4/120, so the departure from a clean square is a part in a hundred at r=0.42r = 0.42 and a part in ten thousand at r=0.1r = 0.1. The fit over the whole range comes out at −2.011, which is the fourth-order term showing up at the L=1L = 1 end and nowhere else.

What an oversampling ratio buys, and at two different rates. computed by solving, not by drawing. A 20 kHz band on a 48 kHz base clock, interpolated by ratios from 1 to 64. The hold's droop at the band edge falls with the SQUARE of the ratio — the fitted exponent over six doublings is -2.0113 — from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it. The nearest image moves out with the FIRST power, exponent 1.0222, from 1.40 times the band edge to 37.4. So one decision buys two things at rates differing by a factor of two in the exponent, and the third quantity — the poles a reconstruction filter needs for sixty decibels — collapses from 20.5 to 3.21 by a ratio of four alone.
Fig. 3 Twice. One doubling has taken the droop from 2.64 dB to 0.629 — a factor of 4.2 — and the image from 1.40 to 3.80 times the band edge, a factor of 2.7. Neither factor is its asymptotic value yet, and both are moving towards it from the same side.

The third quantity is a logarithm and that is why it collapses first

The number of poles is not an independent measurement. It is what the second quantity buys, read through the response of a maximally flat filter: sixty decibels across a frequency ratio ρ\rho needs n=60/(20log10ρ)n = 60/(20\log_{10}\rho) poles.

The band that closes with the order is where that expression comes from and what it is worth. A logarithm of something that rises linearly rises like a logarithm, so the pole count falls like 1/logL1/\log L — which is the slowest of the three and, for exactly that reason, the one that improves most in the first doubling and least afterwards. From the table: ×1 to ×2 takes 20.53 poles to 5.17, a saving of fifteen. ×2 to ×4 takes 5.17 to 3.21, a saving of two. ×16 to ×32 takes 1.91 to 1.60, a saving of a third of a pole, which is not a thing that can be bought.

So the filter order is the reason to oversample by two or four and is not a reason to oversample by sixty-four. Past about eight times the band the analogue filter is already a single pole and there is nothing left in that column to buy. Whatever a ×64 converter is being built for, it is not the reconstruction filter — and the droop column, still falling by four every doubling, is not it either, since 0.0097 dB at ×16 is already two orders below anything measurable.

What is left is the thing this essay has not measured: the quantisation noise, which in a noise-shaped converter falls with the ratio much faster than either of these and is the actual reason the high ratios exist. That is one bit, and where the noise went, and it belongs to a different anchor for a good reason — it is a property of the modulator and not of the hold.

What an oversampling ratio buys, and at two different rates. computed by solving, not by drawing. A 20 kHz band on a 48 kHz base clock, interpolated by ratios from 1 to 64. The hold's droop at the band edge falls with the SQUARE of the ratio — the fitted exponent over six doublings is -2.0113 — from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it. The nearest image moves out with the FIRST power, exponent 1.0222, from 1.40 times the band edge to 37.4. So one decision buys two things at rates differing by a factor of two in the exponent, and the third quantity — the poles a reconstruction filter needs for sixty decibels — collapses from 20.5 to 3.21 by a ratio of four alone.
Fig. 4 Eight times. The droop is 0.0388 dB, which is a fortieth of a Chebyshev’s half-decibel passband ripple; the nearest image is eighteen times the band edge; and the reconstruction filter is 2.38 poles, which in practice is two. Everything the hold contributes to the design has been reduced to nothing by one decision.

The quantity with an exponent of zero

A fourth number belongs on this axis and it does not move at all.

A converter’s aperture error — the amplitude error produced by a clock edge arriving at the wrong time — is the signal’s slew rate times the timing error. For a sinusoid of amplitude AA at frequency ff that is 2πfAσ2\pi f A\sigma, and there is no clock rate in it. Raising the sample rate by sixty-four does not improve it by anything, because the quantity that multiplies the jitter is the frequency of the signal and not the frequency of the clock. A picosecond, read as bits puts a number on that, and the number is the same at every ratio in the table above.

It is worth stating explicitly because it is the one place where the intuition that oversampling “gives more samples and therefore more accuracy” fails outright. More samples of a 20 kHz tone do not make each sample’s timing better; they make more of them, each with the same error, and the errors are not correlated in a way that helps. A converter run at sixty-four times the rate has sixty-four times as many aperture errors in the same second and each one is the same size.

What does improve with the ratio is where the aperture noise lands. The errors are spread across a band that is LL times wider, so the density inside the signal band falls as 1/L1/L — and that is a first power, not the square the droop gets, and it only arrives once a filter has removed everything outside the band. So the honest entry for jitter in this table is: unchanged in total, improved in density by LL, and only after filtering.

Four columns and four exponents: 2-2, +1+1, 1/log-1/\log, and 00. The last is the one that decides whether a converter can be run at the top of its band at all, and it is the only one a ratio cannot touch.

What the ratio costs

Every column improves, so the interesting question is what is on the other side, and there are two things.

The first is clock rate, which is linear in LL and is paid in power and in electromagnetic compliance. A converter clocked at 3.072 MHz radiates at 3.072 MHz, and the image that used to be a filter problem is now a shielding problem — it has not been removed, it has been moved somewhere where a different engineer deals with it.

The second is the interpolating filter itself, whose length grows with the ratio for a fixed transition width. An interpolator that raises the rate by LL must suppress the L1L-1 new images it creates below the old clock, and at a fixed stopband requirement the number of multiplications is roughly proportional to LL — but it runs at LL times the rate, so the arithmetic rate goes as L2L^2. That is the steepest exponent in this essay and it is on the cost side.

Which recovers the usual practice from the arithmetic rather than from convention. Ratios of two to eight are where the analogue side stops being the problem and the digital side has not yet become one; higher ratios are bought for noise shaping and paid for in arithmetic; and the ratio of one is chosen only when there is no digital filter available at all, in which case every number in the first row of the table has to be lived with.

What an oversampling ratio buys, and at two different rates. computed by solving, not by drawing. A 20 kHz band on a 48 kHz base clock, interpolated by ratios from 1 to 64. The hold's droop at the band edge falls with the SQUARE of the ratio — the fitted exponent over six doublings is -2.0113 — from 2.640 dB at the Nyquist rate to 0.0097 at sixteen times it. The nearest image moves out with the FIRST power, exponent 1.0222, from 1.40 times the band edge to 37.4. So one decision buys two things at rates differing by a factor of two in the exponent, and the third quantity — the poles a reconstruction filter needs for sixty decibels — collapses from 20.5 to 3.21 by a ratio of four alone.
Fig. 5 Sixteen. The droop is a hundredth of a decibel and the image is thirty-seven times the band edge, and neither of those numbers is doing any work any more. A quantity that has stopped mattering is as useful a thing to measure as one that has started, because it says where to stop spending.

A worked case, at the two rates the industry actually uses

Put the two common audio clocks side by side, since between them they cover most converters ever built, and the table above is a statement about them rather than about an abstraction.

At 44.1 kHz with a 20 kHz band, rr is 0.4535. The droop at the band edge is 3.17 dB, the nearest image is at 24.1 kHz — 1.21 times the band edge — and a maximally flat filter needs 37 poles for sixty decibels across that transition. Thirty-seven poles is not a filter, it is an admission that the arrangement does not work, and it is the reason no converter has ever been built this way. What is actually built is a converter with an interpolator, which is the ×4 or ×8 row.

At 192 kHz with the same band, rr is 0.1042: the droop is 0.157 dB, the image is at 172 kHz — 8.6 times the band edge — and the filter is 3.21 poles. Every number is comfortable. The clock is four and a third times faster and the design has stopped being about the hold.

The distance between those two rows is the whole practical content of oversampling in this field, and it is two doublings. Everything past it is bought for the modulator, which is a different mechanism in a different part of the converter, and the frequency a sample rate invents is the boundary that both of them are negotiating with.

What the four exponents do not cover

That the exponents are properties of oversampling. They are properties of the two expressions, which are a sinc near its origin and a linear function with an offset. A different reconstruction pulse would give a different first exponent; a different definition of “nearest image” — the centre of the image band rather than its lower edge — would remove the offset and make the second exactly one at every ratio.

That 60 dB is the right stopband requirement. It is a round number chosen so that the third column means something. The pole count is proportional to the requirement, so eighty decibels multiplies every entry by four thirds and changes nothing about the shape of the column.

That a maximally flat filter is what anybody builds. It is the one whose order has a closed form simple enough to sit in a table beside two measurements. An elliptic filter of the same order does far better across the same transition, which lowers every entry in the third column and leaves the argument about its shape — a logarithm of the second quantity — untouched.

That the interpolator’s cost is as precisely known as the three benefits. The L2L^2 above is a scaling argument, not a measurement, and it assumes a single-stage interpolator at a fixed transition width. Multistage interpolation does much better, which is why high ratios are affordable at all, and measuring that properly is a different essay in a different field.

Both exponents fitted, and the one that is only a limit

Both exponents are fitted over the measured columns, not read off the ends, and checked against their integer values — the droop’s over the whole range and the image’s over the upper half.

The asymptotic one is checked to be asymptotic. The full-range fit is required to be measurably larger than the upper-half fit, so that the figure’s own claim about where the power law holds is a measurement rather than a caveat.

Their ratio is checked separately, to within a twentieth, because the argument of the essay is that one exponent is twice the other and that is a statement about the pair rather than about either.

And the pole count is checked to collapse, above twenty at the Nyquist rate and below five at four times it, which is the claim that the third column is spent in the first two doublings.

Three rates, one decision

The structure worth carrying is that a single design parameter can improve several things at once and that the improvements are almost never the same function of it.

A quantity that goes as L2L^{-2}, one that goes as LL, one that goes as 1/logL1/\log L, and a cost that goes as L2L^2. Four different exponents on one knob. The consequence is that there is no single “right” amount of oversampling — there is a right amount for each of the four, and they are two doublings apart from each other.

It also means the usual summary is misleading in a specific way. “Oversampling relaxes the reconstruction filter” is true and is the slowest-improving of the benefits, so a design justified that way is nearly always over-specified: the filter stopped being the binding constraint at ×4 and the ratio was chosen at ×64 for a reason the justification did not mention.

The habit that catches this is putting quantities with different units on one logarithmic axis and fitting their slopes, which is what this figure does and what one dissipation, two exponents does for a quite different pair. A slope is comparable where a value is not, and two slopes a factor of two apart are a different design situation from two slopes that agree.

Still open: the pulse that is not a rectangle, and the ratio the noise wants

A reconstruction pulse with a shape. Every number here is the sinc of a rectangle exactly one clock period long. A converter that emits a shaped pulse — a linear interpolation between samples, which is a triangle and therefore a sinc squared — has twice the droop and twice the image attenuation in decibels, so its two exponents are unchanged and its constants are not. Measuring the pair would say whether the first-order hold is ever the better trade, and at what band edge.

The ratio the noise shaping wants, on the same axes. The three columns here are all properties of the hold. A noise-shaped converter’s in-band quantisation noise falls as L(2m+1)L^{-(2m+1)} for an mm-th-order modulator, which is the steepest exponent in the whole subject and is what actually sets the ratios in use. Drawn on this figure’s axes it would dwarf all three columns and would show, immediately, which column any given converter’s ratio was chosen for.

And the arithmetic cost measured rather than argued. The L2L^2 above is a scaling argument. A multistage interpolator’s actual multiplication rate against LL, for a fixed stopband, is a measurement that could sit on the same axes as its benefit — and it is the only way to say where the total stops improving.

Part 4 on reconstruction

One argument about Reconstruction, and one of 4 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Anti imaging filterDesign tradeoffFilter orderOversamplingPower law fitReconstructionZero-order hold