Where a signal becomes a number

The pole a straight line is worth

A converter that joins its samples with straight lines instead of holding each one is convolving with a triangle rather than a rectangle, and its transform is the hold's sinc squared. Every decibel doubles: 5.28 dB of droop at a 20 kHz band edge on a 48 kHz clock instead of 2.64, and 5.85 dB of image rejection instead of 2.92. Both exponents of the oversampling ratio are unchanged. The second sinc is worth exactly one pole of reconstruction filter at every ratio, because a sinc's rejection of the nearest image is twenty times the log of that image's distance from the band edge, identically. A single analogue pole allowed the same droop gives a fraction of that rejection. The price is a whole clock of delay rather than half.

Assumes: The staircase on the way out · The frequency a sample rate invents

The staircase on the way out established what a converter’s output is. It is not a train of impulses but a staircase, each sample held for a whole clock period, and holding is a convolution with a rectangle, so the spectrum is multiplied by a sinc. The nulls are where nothing is followed that sinc out to the images and found it rejects each image at the image’s own frequency. One knob, and the two exponents it turns found that oversampling shrinks the droop as the square of the ratio while moving the images out only as its first power.

Every one of those numbers belongs to a rectangle one clock long. That essay ended by asking about a pulse with a shape. The simplest one after the rectangle is the line. A converter, or the interpolator in front of it, can draw a straight line from each sample to the next instead of holding the first until the second arrives. This page asks what that line does to both exponents and their constants, and whether it is ever the better trade.

A triangle is a rectangle twice

Joining samples with straight lines is a convolution too. Each sample becomes a triangle two clock periods wide, rising from zero at the previous sample’s time to its full value at its own and back to zero at the next, and the sum of those overlapping triangles is exactly the broken line through the samples. A triangle of that width is a one-clock rectangle convolved with itself. Convolution in time is multiplication in frequency, so the triangle’s transform is the rectangle’s sinc multiplied by itself.

Joining the samples with straight lines squares the hold's sinc: 5.28 dB of droop and 5.8 dB of image rejection, both doubledcomputed by solving, not by drawing. A 20 kHz tone on a 48 kHz clock, reconstructed two ways: held for a whole clock period (dashed envelope, open stems) and joined sample to sample by straight lines (solid envelope, filled stems). Every line is read from a transform of that waveform itself. The held one comes out 2.640 dB down and its nearest image 2.92 dB below it; the interpolated one 5.280 dB down with the image 5.85 dB below. A straight line between samples is a triangle two clocks wide, which is a rectangle convolved with itself, so its transform is the sinc squared: every decibel doubles and the nulls stay where they were.-60-40-20000.50011.5022.503frequency, as a multiple of the sample rateamplitude (dB, full scale)dashed: heldthe sincsolid: interpolatedthe sinc squaredsignal20 kHz = 0.417 fsdroop, held2.640 dB…interpolated5.280 dBfirst image, held−2.92 dB…interpolated−5.85 dBtwo routes agree0.059%solved, then checked — two waveforms, each transformedevery decibel twice
Fig. 1 A 20 kHz tone on a 48 kHz clock reconstructed two ways: held (dashed envelope) and joined sample to sample by straight lines (solid envelope), with every line read from a transform of that waveform itself. Held, the tone comes out 2.640 dB down and its nearest image 2.92 dB below it; interpolated, 5.280 dB down and 5.85 dB below. Every line of the interpolated waveform is twice the held one’s in decibels, and the nulls have not moved.

The figure measures that on both waveforms separately rather than assuming it. Each is built sample by sample as a converter would build it, and every line is read from that record’s own transform. The interpolated waveform’s lines agree with the squared sinc to 0.06 per cent, and each is exactly twice the held waveform’s line in decibels. At 20 kHz on a 48 kHz clock the held tone comes out 2.64 dB down and the interpolated one 5.28 dB. The nearest image, at 28 kHz, sits 2.92 dB below the held tone and 5.85 dB below the interpolated one.

Two things follow at once. The droop doubles, so whatever the rectangle’s droop cost, the triangle costs twice. And the image rejection doubles, since a ratio of squared values is the square of the ratio. The nulls are the same nulls, at every multiple of the clock, because squaring a zero leaves it a zero. Nothing about where the sinc puts things changes, only how deep they are.

Joining the samples with straight lines squares the hold's sinc: 0.80 dB of droop and 28.0 dB of image rejection, both doubled. computed by solving, not by drawing. A 8 kHz tone on a 48 kHz clock, reconstructed two ways: held for a whole clock period (dashed envelope, open stems) and joined sample to sample by straight lines (solid envelope, filled stems). Every line is read from a transform of that waveform itself. The held one comes out 0.401 dB down and its nearest image 13.98 dB below it; the interpolated one 0.801 dB down with the image 27.96 dB below. A straight line between samples is a triangle two clocks wide, which is a rectangle convolved with itself, so its transform is the sinc squared: every decibel doubles and the nulls stay where they were.
Fig. 2 The same comparison with the tone at 8 kHz, a sixth of the clock: the held tone is down 0.401 dB and the interpolated one 0.801 dB, while the nearest image falls from 13.98 dB below the tone to 27.96 dB below it.

Lower in the band the same doubling is less alarming and more useful. At 8 kHz the held tone droops by 0.40 dB and the interpolated one by 0.80, a cost of four tenths of a decibel, while the nearest image goes from 14.0 dB below the tone to 28.0. At 4 kHz the droop goes from 0.10 dB to 0.20 and the image from 20.8 dB down to 41.7. The droop is small where the signal is low and the image rejection is large there, and squaring multiplies both by the same factor of two. So the doubling looks very different depending on which of the two is the binding constraint.

Both exponents, unchanged

The earlier essay measured two exponents of the oversampling ratio: the droop at a 20 kHz band edge falling as its square, and the image moving out as its first power. Doubling a number in decibels is multiplying it by a constant, which on logarithmic axes is a constant offset, and a constant offset changes no slope.

The interpolated hold keeps both exponents and doubles both constants. computed by solving, not by drawing. The droop at a 20 kHz band edge and the rejection of the nearest image, for a clock of 48 kHz times the oversampling ratio, held (dashed) and interpolated (solid). On a log axis a doubling in decibels is a constant offset, so the droops fall with the same fitted exponent, −2.011, from 2.640 and 5.280 dB at the Nyquist rate to 0.00061 and 0.00121 dB at sixty-four times it. The rejections grow by 6.16 and 12.31 dB per doubling, from 2.92 and 5.85 dB to 43.7 and 87.3 dB.
Fig. 3 The droop at a 20 kHz band edge and the nearest image’s rejection against the oversampling ratio, held (dashed) and interpolated (solid). The droops fall with one fitted exponent, −2.011, from 2.640 and 5.280 dB at the Nyquist rate to 0.0006 and 0.0012 dB at sixty-four times it. The rejections grow 6.16 and 12.31 dB per doubling, from 2.92 and 5.85 dB to 43.7 and 87.3.

The droop falls with a fitted exponent of −2.011 for both holds, from 2.64 and 5.28 dB at the Nyquist rate to six and twelve ten-thousandths of a decibel at sixty-four times it. The rejection of the nearest image grows about six decibels per doubling of the clock for the held waveform and exactly twice that for the interpolated one. So the answer to the first half of the question left open is that the exponents are the rectangle’s and the constants are doubled, as expected. The interesting part is what the doubled constant is worth in the currency a designer spends, and that turns out to be exact.

Exactly one pole

The sinc’s rejection of the nearest image has a closed form that the earlier essays measured without writing down. With the band edge at a fraction xx of the clock, the nearest image is at 1x1 - x, and the sinc at those two points is sinπx/πx\sin \pi x / \pi x and sinπ(1x)/π(1x)\sin \pi(1-x) / \pi(1-x). The two sines are equal. So the ratio of the signal to its image is (1x)/x(1-x)/x, which is the image’s frequency over the band edge’s, and the rejection in decibels is twenty times the log of the image’s distance from the band edge, exactly, at every ratio.

That is also exactly what one pole of a reconstruction filter gives over the same distance, at twenty decibels a decade. So a hold is worth one pole, and a second sinc is worth a second.

The second sinc is worth exactly one pole of reconstruction filter, at every oversampling ratio. computed by solving, not by drawing. The order of maximally flat reconstruction filter needed for sixty decibels at the nearest image, counting the hold's own rejection there, held (dashed) and interpolated (solid), for a 20 kHz band on 48 kHz times the ratio. At the Nyquist rate the two need 19.53 and 18.53 poles, exactly one apart, and they are exactly one apart at every ratio: the rejection a sinc gives the nearest image, relative to the signal, is 20 log of the image's distance from the band edge, which is what one pole gives over that distance. At 16 times the interpolated hold needs no filter at all for sixty decibels and the held one 0.91 poles.
Fig. 4 The order of maximally flat reconstruction filter needed for sixty decibels at the nearest image after each hold’s own rejection, held (dashed) and interpolated (solid). At the Nyquist rate 19.53 against 18.53 poles, and exactly one apart at every ratio. At sixteen times the interpolated hold needs no filter for sixty decibels and the held one 0.91 poles.

Counting the hold’s own share, a held converter at the Nyquist rate needs 19.53 poles for sixty decibels at its nearest image, and an interpolating one 18.53. One knob, and the two exponents it turns quoted 20.53 for the same requirement because it asked the filter for all sixty decibels, which is the zero-order hold’s pole left uncounted. The difference between the two holds is one pole to twelve decimal places at every ratio, and at four times oversampling it is 2.21 poles against 1.21. At sixteen times the interpolated hold reaches sixty decibels with no filter at all, while the held one still needs most of a pole.

The one-pole result is the useful form of the answer. Linear interpolation is a first-order section of reconstruction filter that lives in the converter, and it costs neither a component nor a stage of analogue noise.

Where in the band the pole is worth most

The closed form says more than the one-pole result. A tone at a fraction xx of the clock has its nearest image rejected by 20log10((1x)/x)20\log_{10}((1-x)/x) per sinc, so the rejection is large low in the band and shrinks towards the top of it. At a twelfth of the clock one sinc gives 20.8 dB and two give 41.7; at a third of the clock one sinc gives 6.0 dB and two give 12.0. A reconstruction filter has to meet its requirement at the band edge, where the rejection is least, so the band edge’s number is the one that sets the filter, and the one-pole result is stated there.

At half the clock both numbers are zero. The tone and its image are at the same frequency and the same height, and squaring zero decibels still gives zero. The nulls are where nothing is found the held converter giving nothing to a tone at half the clock, and the interpolated one gives it exactly as much. No hold, of any order, can separate a tone from an image that coincides with it. That is the boundary the frequency a sample rate invents found on the way in, met again on the way out, and oversampling is the only thing that moves the band edge away from it.

The asymmetry between the two ends of a converter is worth noticing. On the way in, what the filter in front costs found the anti-aliasing filter’s order set by how close the folding frequency sits to the band, with nothing on the converter’s side to share the work. On the way out the hold shares it, and the choice of hold decides how large a share it takes: one pole for a rectangle, two for a triangle.

The cheaper pole

A pole is not free even when it is an analogue one, and the real comparison is what each costs in droop. The second sinc costs one more helping of the hold’s own droop: 2.64 dB at the Nyquist rate, 0.156 at four times, falling as the square of the ratio. An analogue pole is a first-order low-pass and its cost depends on where its corner goes.

For the droop the second sinc costs, an analogue pole gives the image 5.5 dB where the sinc gives 18.7. computed by solving, not by drawing. What the nearest image is given, relative to the signal, three ways, for a 20 kHz band on 48 kHz times the ratio: by the interpolated hold's second sinc (solid), by a single analogue pole whose corner is placed so that it costs exactly the same band-edge droop as that sinc (dashed), and by a pole with its corner on the band edge, which costs 3.01 dB there (faint). At four times the sinc gives 18.69 dB for 0.156 dB of droop; the pole allowed the same droop gives 5.52 dB, and the pole at the band edge 15.74 dB for 19 times the droop. A pole reaches the sinc's rejection only as its corner goes to zero, and the sinc has it for the price of one more helping of the hold's own droop.
Fig. 5 The nearest image’s rejection relative to the signal, bought three ways: by the second sinc (solid), by a single analogue pole allowed exactly the second sinc’s band-edge droop (dashed), and by a pole with its corner on the band edge, paying 3.01 dB there (faint). At four times the sinc gives 18.7 dB for 0.156 dB of droop, and the pole given the same droop gives 5.5.

Allowed exactly the droop the second sinc costs, an analogue pole must put its corner well above the band, and it then gives the nearest image far less. At four times oversampling the sinc gives 18.7 dB for 0.156 dB of droop and the pole given the same droop gives 5.5 dB. At sixteen times it is 31.5 against 6.1. Even at the Nyquist rate, where the sinc’s droop is largest, the sinc gives 2.9 dB for 2.64 dB of droop and the pole 1.6. Placed with its corner on the band edge, the pole pays 3.01 dB there, twenty times the sinc’s price at four times oversampling, and still gives three decibels less than the sinc.

A single pole can only match the sinc’s rejection as its corner goes to zero, which means unbounded droop. The sinc gets the full asymptotic value at once because its magnitude near a null falls like a single zero at the clock frequency, and a zero pulls the image down harder than a pole far below it does. So a first-order hold is not merely a pole’s worth of rejection. It is the cheapest pole available, cheaper in droop than any analogue pole at every oversampling ratio.

Flatness, and the two currencies it is bought in put a price on the droop: corrected digitally it costs exactly its own size in headroom, and it costs no signal-to-noise ratio at all, because the droop takes the signal and everything with it down together. The same applies here. The interpolated hold’s extra droop is a known curve that a digital pre-emphasis can take out at the price of that many decibels of headroom, and at four times oversampling that is 0.16 dB.

A clock of waiting

The magnitude does not show the cost that makes the first-order hold rare in practice. The straight line from one sample to the next can only be drawn once the next sample exists.

Held, the output is half a clock late; joined by straight lines, a whole clock late. computed by solving, not by drawing. A 4 kHz tone sampled at 48 kHz (dots), the input (faint), held for each clock (dashed) and interpolated causally from each sample to the next (solid). The fundamental of the held waveform lags the input by 0.500 of a clock, 10.42 µs; the interpolated one by 1.000, 20.83 µs. The straight line from one sample to the next can only be drawn once the next has arrived, and that clock of waiting is the price the magnitude response does not show.
Fig. 6 A 4 kHz tone sampled at 48 kHz (dots), the input (faint), the held waveform (dashed) and the causally interpolated one (solid). The held waveform’s fundamental lags the input by half a clock, 10.42 µs, and the interpolated one’s by a whole clock, 20.83 µs.

A held output lags by half a clock, since a rectangle’s centre is half a period after its start. A causal linear interpolation lags by a whole clock, since during each period it is drawing the line towards a sample that arrived at the start of that period, from the one before. At 48 kHz that is 20.8 µs against 10.4, measured from the phase of each waveform’s fundamental. It is a pure delay, the same at every frequency, so it distorts nothing, which is the property flat delay, bought with more delay had to build an all-pass to obtain. For a converter feeding a loudspeaker a further ten microseconds is nothing. For a converter inside a control loop it is phase lag that grows with frequency, and a loop crossing over at a tenth of the clock loses eighteen degrees of margin to the extra half clock alone.

The delay deserves one more comparison, because the analogue pole the second sinc replaces is not free of delay either. A first-order low-pass delays low frequencies by 1/(2πfc)1/(2\pi f_c), which for a corner on a 20 kHz band edge is 7.96 µs. The interpolated hold’s extra half clock is 10.4 µs at the Nyquist rate, a little more than the pole it replaces, and 2.6 µs at four times oversampling, a third of it. So above the lowest ratios the straight line is not only the cheaper pole in droop but also the shorter one in delay, and the delay argument against it holds only where the clock sits close to the band.

What a designer should take

Linear interpolation squares the hold’s response. It doubles the droop and the image rejection in decibels and changes neither exponent of the oversampling ratio. The rejection it adds is worth exactly one pole of reconstruction filter at every ratio, and in droop it is cheaper than any analogue pole could be. What it costs is half a clock of extra delay. So the first-order hold is the better trade wherever the extra delay does not matter, at every band edge, and it is most valuable at modest oversampling, where one pole is a large fraction of what the filter needs. At four times it halves the filter from 2.21 poles to 1.21, and at sixteen it removes the filter.

How the numbers were obtained

The two waveforms are built at 256 points per clock from a tone whose frequency is a whole number of cycles in a whole number of clocks, twelve clocks here, so every spectral line lands on a bin and no window is needed. The held waveform takes each sample’s value for one clock; the interpolated one runs during each clock from the previous sample to the current one. Each record is transformed and its lines are compared with the sinc and the squared sinc at their own frequencies. The droops and rejections against the ratio are the closed forms, which those transforms confirm. The filter orders use the same maximally flat rule as the earlier essay, sixty decibels over twenty decibels a decade times the log of the image’s distance, after subtracting each hold’s own rejection there. The analogue pole’s corner is solved from the droop it is allowed. The delays are read from the phase of each waveform’s fundamental over two whole periods, by midpoint sums on 24,000 points.

What it leaves out

Higher-order interpolation. A quadratic or cubic spline is a longer pulse, three or four rectangles convolved, and raises the sinc to the third or fourth power. The one-pole law generalises to one pole per convolution with the rectangle, but the delay grows with each, and there is a point at which the delay is a larger cost than the filter it replaces.

Interpolators that are not holds at all. A converter behind a digital interpolation filter raises the rate first and holds afterwards, and its images are then set by the digital filter rather than by any sinc. The first-order hold applies to the last stage, at whatever rate reaches it.

And the droop’s correction as a filter of its own. The pre-emphasis that flattens the squared sinc is steeper than the one that flattens the sinc, and near half the clock it asks for gain the headroom cannot give.

Still open: the cubic pulse, the delay in a loop, and the correction at the Nyquist rate

The cubic pulse. A cubic interpolator is four rectangles convolved together and should be worth three poles over the held converter, at a delay of two clocks. Measuring its images on its own waveform would test the one-pole-per-rectangle law, and setting the poles saved against the delay added would say where on the oversampling axis a longer pulse stops paying.

The extra half clock inside a loop. A converter in a digitally controlled loop pays the hold’s delay as phase lag. The half clock that linear interpolation adds costs a fixed number of degrees at the crossover, and there is a crossover, as a fraction of the clock, at which the margin lost outweighs the filter poles saved. Finding it would say which control loops can use the cheaper pole.

The correction at the Nyquist rate. At the Nyquist rate the interpolated hold droops 5.28 dB at a 20 kHz band edge and 7.84 dB at half the clock. A digital pre-emphasis for that is an inverse squared sinc, and the question is whether a short filter can do it to a tenth of a decibel without lifting the images it was meant to leave alone.

Part 5 on reconstruction

One argument about Reconstruction, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Anti imaging filterDesign tradeoffFilter orderGroup delayOversamplingReconstructionZero-order hold