One knob, and the two exponents it turns
Assumes: The staircase on the way out · The frequency a sample rate invents
Three earlier essays here have measured consequences of one rectangle and every one of them has ended at the same place: the answer is a function of , the band edge as a fraction of the clock, and is the one quantity a designer controls.
The staircase on the way out found a droop of . The nulls are where nothing is found an image rejection of and a filter transition ratio of . Flatness, and the two currencies it is bought in found two prices for correcting the first, both functions of the same .
Lowering is oversampling, and it is normally described as buying a gentler filter. It buys at least three separate things, and they improve at three different rates. The rates are what decide which of them a given ratio is actually being bought for, and none of the three is usually measured against the others on one axis.
Three quantities, and their exponents fitted over six doublings
The arrangement is the ordinary one. A signal band that stops at 20 kHz, a base clock of 48 kHz, and an interpolator that raises the sample rate by before the converter. Raising the rate does not change the band and does not change the samples’ information; it changes where everything sits relative to the clock.
| ratio | clock | droop at 20 kHz | nearest image | poles for 60 dB |
|---|---|---|---|---|
| ×1 | 48 kHz | 2.640 dB | 1.40× the band | 20.53 |
| ×2 | 96 kHz | 0.6292 dB | 3.80× | 5.17 |
| ×4 | 192 kHz | 0.1556 dB | 8.60× | 3.21 |
| ×8 | 384 kHz | 0.0388 dB | 18.20× | 2.38 |
| ×16 | 768 kHz | 0.0097 dB | 37.40× | 1.91 |
| ×32 | 1.536 MHz | 0.0024 dB | 75.80× | 1.60 |
| ×64 | 3.072 MHz | 0.0006 dB | 152.6× | 1.37 |
Read down the droop column: each doubling divides it by four. Read down the image column: each doubling roughly doubles it. Read down the third: each doubling takes about a fifth off it, and the fifth gets smaller.
Fitted over the whole range rather than read off two points — a fit over two points is an arithmetic identity and proves nothing — the exponents come out at −2.011 for the droop and 1.022 for the image separation. Both are checked against their integer values, and the second is checked over the upper half of the range for a reason given below.
The origin of the first is a series. Near the origin the sinc is , so the decibels are to leading order, and is inversely proportional to . The droop is therefore quadratic in the ratio and there is nothing to choose about it.
The origin of the second is subtraction. The nearest image sits at , so its distance from the band edge is — proportional to with an offset. The offset is not negligible at small : at the Nyquist rate the image is at 1.40 times the band edge rather than at 2.40, because the subtraction has taken away most of it.
The exponent that is asymptotic, and where it says so
The first power is a limit rather than a law, and the rule here is to say where a law stops being one.
Over the full range from one to sixty-four the fitted exponent for the image distance is 1.109, not 1.0. Over the upper half — eight to sixty-four — it is 1.022. The difference is eleven per cent and it is entirely the in , which matters when is small and does not when it is large.
Eleven per cent would be a rounding error if the small- end were an academic corner. It is not: a converter with no interpolator at all sits at , and a great many instrument outputs, arbitrary waveform generators and control DACs do exactly that. So the region where the first power is wrong by eleven per cent is precisely the region a designer without an interpolating filter is working in, and the correct statement there is the subtraction rather than the power law.
The droop’s exponent has the same structure and is better behaved. Its series is in and the next term is , so the departure from a clean square is a part in a hundred at and a part in ten thousand at . The fit over the whole range comes out at −2.011, which is the fourth-order term showing up at the end and nowhere else.
The third quantity is a logarithm and that is why it collapses first
The number of poles is not an independent measurement. It is what the second quantity buys, read through the response of a maximally flat filter: sixty decibels across a frequency ratio needs poles.
The band that closes with the order is where that expression comes from and what it is worth. A logarithm of something that rises linearly rises like a logarithm, so the pole count falls like — which is the slowest of the three and, for exactly that reason, the one that improves most in the first doubling and least afterwards. From the table: ×1 to ×2 takes 20.53 poles to 5.17, a saving of fifteen. ×2 to ×4 takes 5.17 to 3.21, a saving of two. ×16 to ×32 takes 1.91 to 1.60, a saving of a third of a pole, which is not a thing that can be bought.
So the filter order is the reason to oversample by two or four and is not a reason to oversample by sixty-four. Past about eight times the band the analogue filter is already a single pole and there is nothing left in that column to buy. Whatever a ×64 converter is being built for, it is not the reconstruction filter — and the droop column, still falling by four every doubling, is not it either, since 0.0097 dB at ×16 is already two orders below anything measurable.
What is left is the thing this essay has not measured: the quantisation noise, which in a noise-shaped converter falls with the ratio much faster than either of these and is the actual reason the high ratios exist. That is one bit, and where the noise went, and it belongs to a different anchor for a good reason — it is a property of the modulator and not of the hold.
The quantity with an exponent of zero
A fourth number belongs on this axis and it does not move at all.
A converter’s aperture error — the amplitude error produced by a clock edge arriving at the wrong time — is the signal’s slew rate times the timing error. For a sinusoid of amplitude at frequency that is , and there is no clock rate in it. Raising the sample rate by sixty-four does not improve it by anything, because the quantity that multiplies the jitter is the frequency of the signal and not the frequency of the clock. A picosecond, read as bits puts a number on that, and the number is the same at every ratio in the table above.
It is worth stating explicitly because it is the one place where the intuition that oversampling “gives more samples and therefore more accuracy” fails outright. More samples of a 20 kHz tone do not make each sample’s timing better; they make more of them, each with the same error, and the errors are not correlated in a way that helps. A converter run at sixty-four times the rate has sixty-four times as many aperture errors in the same second and each one is the same size.
What does improve with the ratio is where the aperture noise lands. The errors are spread across a band that is times wider, so the density inside the signal band falls as — and that is a first power, not the square the droop gets, and it only arrives once a filter has removed everything outside the band. So the honest entry for jitter in this table is: unchanged in total, improved in density by , and only after filtering.
Four columns and four exponents: , , , and . The last is the one that decides whether a converter can be run at the top of its band at all, and it is the only one a ratio cannot touch.
What the ratio costs
Every column improves, so the interesting question is what is on the other side, and there are two things.
The first is clock rate, which is linear in and is paid in power and in electromagnetic compliance. A converter clocked at 3.072 MHz radiates at 3.072 MHz, and the image that used to be a filter problem is now a shielding problem — it has not been removed, it has been moved somewhere where a different engineer deals with it.
The second is the interpolating filter itself, whose length grows with the ratio for a fixed transition width. An interpolator that raises the rate by must suppress the new images it creates below the old clock, and at a fixed stopband requirement the number of multiplications is roughly proportional to — but it runs at times the rate, so the arithmetic rate goes as . That is the steepest exponent in this essay and it is on the cost side.
Which recovers the usual practice from the arithmetic rather than from convention. Ratios of two to eight are where the analogue side stops being the problem and the digital side has not yet become one; higher ratios are bought for noise shaping and paid for in arithmetic; and the ratio of one is chosen only when there is no digital filter available at all, in which case every number in the first row of the table has to be lived with.
A worked case, at the two rates the industry actually uses
Put the two common audio clocks side by side, since between them they cover most converters ever built, and the table above is a statement about them rather than about an abstraction.
At 44.1 kHz with a 20 kHz band, is 0.4535. The droop at the band edge is 3.17 dB, the nearest image is at 24.1 kHz — 1.21 times the band edge — and a maximally flat filter needs 37 poles for sixty decibels across that transition. Thirty-seven poles is not a filter, it is an admission that the arrangement does not work, and it is the reason no converter has ever been built this way. What is actually built is a converter with an interpolator, which is the ×4 or ×8 row.
At 192 kHz with the same band, is 0.1042: the droop is 0.157 dB, the image is at 172 kHz — 8.6 times the band edge — and the filter is 3.21 poles. Every number is comfortable. The clock is four and a third times faster and the design has stopped being about the hold.
The distance between those two rows is the whole practical content of oversampling in this field, and it is two doublings. Everything past it is bought for the modulator, which is a different mechanism in a different part of the converter, and the frequency a sample rate invents is the boundary that both of them are negotiating with.
What the four exponents do not cover
That the exponents are properties of oversampling. They are properties of the two expressions, which are a sinc near its origin and a linear function with an offset. A different reconstruction pulse would give a different first exponent; a different definition of “nearest image” — the centre of the image band rather than its lower edge — would remove the offset and make the second exactly one at every ratio.
That 60 dB is the right stopband requirement. It is a round number chosen so that the third column means something. The pole count is proportional to the requirement, so eighty decibels multiplies every entry by four thirds and changes nothing about the shape of the column.
That a maximally flat filter is what anybody builds. It is the one whose order has a closed form simple enough to sit in a table beside two measurements. An elliptic filter of the same order does far better across the same transition, which lowers every entry in the third column and leaves the argument about its shape — a logarithm of the second quantity — untouched.
That the interpolator’s cost is as precisely known as the three benefits. The above is a scaling argument, not a measurement, and it assumes a single-stage interpolator at a fixed transition width. Multistage interpolation does much better, which is why high ratios are affordable at all, and measuring that properly is a different essay in a different field.
Both exponents fitted, and the one that is only a limit
Both exponents are fitted over the measured columns, not read off the ends, and checked against their integer values — the droop’s over the whole range and the image’s over the upper half.
The asymptotic one is checked to be asymptotic. The full-range fit is required to be measurably larger than the upper-half fit, so that the figure’s own claim about where the power law holds is a measurement rather than a caveat.
Their ratio is checked separately, to within a twentieth, because the argument of the essay is that one exponent is twice the other and that is a statement about the pair rather than about either.
And the pole count is checked to collapse, above twenty at the Nyquist rate and below five at four times it, which is the claim that the third column is spent in the first two doublings.
Three rates, one decision
The structure worth carrying is that a single design parameter can improve several things at once and that the improvements are almost never the same function of it.
A quantity that goes as , one that goes as , one that goes as , and a cost that goes as . Four different exponents on one knob. The consequence is that there is no single “right” amount of oversampling — there is a right amount for each of the four, and they are two doublings apart from each other.
It also means the usual summary is misleading in a specific way. “Oversampling relaxes the reconstruction filter” is true and is the slowest-improving of the benefits, so a design justified that way is nearly always over-specified: the filter stopped being the binding constraint at ×4 and the ratio was chosen at ×64 for a reason the justification did not mention.
The habit that catches this is putting quantities with different units on one logarithmic axis and fitting their slopes, which is what this figure does and what one dissipation, two exponents does for a quite different pair. A slope is comparable where a value is not, and two slopes a factor of two apart are a different design situation from two slopes that agree.
Still open: the pulse that is not a rectangle, and the ratio the noise wants
A reconstruction pulse with a shape. Every number here is the sinc of a rectangle exactly one clock period long. A converter that emits a shaped pulse — a linear interpolation between samples, which is a triangle and therefore a sinc squared — has twice the droop and twice the image attenuation in decibels, so its two exponents are unchanged and its constants are not. Measuring the pair would say whether the first-order hold is ever the better trade, and at what band edge.
The ratio the noise shaping wants, on the same axes. The three columns here are all properties of the hold. A noise-shaped converter’s in-band quantisation noise falls as for an -th-order modulator, which is the steepest exponent in the whole subject and is what actually sets the ratios in use. Drawn on this figure’s axes it would dwarf all three columns and would show, immediately, which column any given converter’s ratio was chosen for.
And the arithmetic cost measured rather than argued. The above is a scaling argument. A multistage interpolator’s actual multiplication rate against , for a fixed stopband, is a measurement that could sit on the same axes as its benefit — and it is the only way to say where the total stops improving.
Part 4 on reconstruction
One argument about Reconstruction, and one of 4 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:
The objects named here
The third axis, after the field and the idea: the things themselves, and every essay that touches each one.
Anti imaging filterDesign tradeoffFilter orderOversamplingPower law fitReconstructionZero-order hold
- The digits the arithmetic did not have design tradeoff, filter order
- The factor the expression leaves out design tradeoff, power law fit
- The loop that is worse at full scale design tradeoff, oversampling
- The product that is not the third design tradeoff, power law fit
- The selectivity that is not free design tradeoff, filter order
- Two loops, and the mismatch between them design tradeoff, oversampling