Where a signal becomes a number

The frequency a sample rate invents

Every other boundary on this site is a model getting gradually worse. This one has no gradient at all: below half the sample rate a set of samples has one sinusoid through it, above half the sample rate it has another, and the two sets of numbers are identical to three parts in ten thousand billion. Nothing is attenuated, nothing is distorted, and there is no measurement of the samples that could say which signal was there.

Assumes: One solve, read four ways · Every model has an edge

Every boundary this collection has drawn so far degrades. An ideal amplifier is one per cent low at 1.4 kHz and worse at 2 kHz; a capacitor is ten per cent off at 4.69 MHz and further off above it; Kirchhoff’s laws are one degree out at 3.97 MHz on ten centimetres of track and two degrees out somewhere higher. In every case the model gets steadily less true, the error is a continuous function of the thing that broke it, and a reader who knows the number knows how much margin they have.

A sample rate is not like that. Below half of it a set of samples determines one signal. Above half of it the same set of samples determines a different signal, at full amplitude, with nothing in the numbers to indicate that anything has happened. There is no gradual region, no margin, and no measurement of the samples that can tell the two apart — because there is nothing to tell apart. They are the same numbers.

A 9.0 kHz input sampled at 10 kHz arrives as 1.0 kHzcomputed by solving, not by drawing. The dots are the samples. The input at 9.00 kHz is above half the 10 kHz rate, and every dot also lies on the 1.00 kHz curve drawn beside it — the two sample sequences differ by 9.3e-15, which is the arithmetic and not a small effect. Nothing is attenuated and nothing is distorted: the samples are the samples of a different signal, at full amplitude, and there is no measurement of them that could say which one was there.the input, 9.00 kHzwhat the samples say, 1.00 kHzsample rate10.0 kHzhalf of it5.00 kHzinput9.00 kHzreported as1.00 kHzsamples differ by9.3e-15solved, then checked — one sequence, two sourceshalf the sample rate: 5.00 kHz
Fig. 1 Two sinusoids and one set of samples. The dots are what a converter running at 10 kHz hands on when it is given 9 kHz; the second curve is 1 kHz with its phase reversed, and every dot lies on it as well. The largest difference between the two sample sequences is at the arithmetic’s floor. The slider walks the input across half the rate; below it the second curve is the first one.

What is being claimed, and how it is checked

The usual statement is that frequencies above half the sample rate “fold back” or “appear as” lower ones. Both phrasings suggest a process — something happening to the signal on the way in — and it is worth being exact, because the exactness is the whole content.

Sampling a sinusoid of frequency f at rate fₛ produces the sequence

x[k]=Asin ⁣(2πfkfs+ϕ)x[k] = A\sin\!\left(2\pi f \frac{k}{f_\mathrm{s}} + \phi\right)

and the sine’s argument is only ever evaluated at multiples of 1/fₛ. Adding a whole multiple of fₛ to f adds a whole multiple of 2πk to the argument, which changes nothing. Subtracting f from fₛ negates the argument, which flips the sign. So for every input there is an infinite family of other inputs — fₛ − f, fₛ + f, 2fₛ − f, 2fₛ + f, and onwards — whose samples are the same numbers.

That is an algebraic identity and this site does not draw algebraic identities. The check in scripts/netcheck.mjs generates both sequences independently, from separate calls with separate frequencies, and differences them. Across four sample rates — 8 kHz, 10 kHz, 44.1 kHz and 1 MHz — four input fractions of each, and three image orders, that is ninety-six pairs of ninety-six-sample sequences, and the largest disagreement anywhere in them is 3.7 × 10⁻¹³. Which is to say: not close, not similar, not within tolerance. The same.

The phase reversal on the odd images is the part that is easy to lose, and the gate holds it separately for a reason. A first version of the figure drew fₛ − f without the reversal, the dots sat on neither curve, and the drawing read as a defect in the sampler rather than as a missing minus sign in the identity being illustrated. The check now generates that mistake deliberately and requires it to fail: without the reversal the two sequences differ by 1.90 on an amplitude of 1, which is nearly twice the signal — a large, obvious, wrong answer, and exactly the kind that a tolerance widened by half a decibel would have absorbed.

Why “half the sample rate” and not “the sample rate”

The images arrive in pairs — fₛ − f below the rate and fₛ + f above it — so the axis folds about fₛ/2 rather than about fₛ. An input at 4 kHz sampled at 10 kHz has its first image at 6 kHz; an input at 6 kHz has its first image at 4 kHz. The two are the same statement read from either end, and the frequency at which an input and its own first image coincide is 5 kHz, which is half the rate.

So the recoverable band is not “up to the sample rate, with some care above half of it”. It is up to half the rate, full stop, and the reason the boundary is sharp is that it is a coincidence rather than a degradation: at 4.9 kHz the input and its image are 200 Hz apart and both are inside the band; at 5.1 kHz they have swapped, and the one inside the band is the wrong one.

A 4.5 kHz input sampled at 10 kHz arrives as itself. computed by solving, not by drawing. The dots are the samples. The input at 4.50 kHz is below half the 10 kHz rate, so the only sinusoid through these dots below half the rate is the one that was sampled. That is the whole of what the sampling theorem promises, and it stops promising it at 5.00 kHz.
Fig. 2 The same generator a little below the boundary. The image is at 5.5 kHz and is outside the band, so the only sinusoid through these dots that is inside the band is the one that was sampled. Half a kilohertz higher the sentence reverses and nothing about the dots changes to say so.

This is also why the boundary is quoted as a property of the rate rather than of the signal. The converter has no view about which of the family it was given. It reports a set of numbers, and the statement “these are samples of 1 kHz” is a statement somebody downstream makes, on the strength of an assumption nobody checked.

What it costs, next to the site’s other boundaries

Put on the axis the collection already uses for this, the sample rate is a boundary of a different species. Four of the site’s boundaries fit on one frequency axis for a ten-centimetre board, and the ordering across them is itself a measured result rather than a general truth.

Where four of this site's models stop being true. In order: the ideal operational amplifier at 1.42 kHz, a 10 V output at full amplitude at 7.96 kHz, Kirchhoff's laws on 10.0 cm at 3.97 MHz, the ideal 100 nF capacitor at 4.69 MHz. The fifth boundary is an amplitude rather than a frequency and cannot share this axis: a small-signal model is 1% wrong above 7.3 mV, at every frequency there is.
Fig. 3 The four boundaries this site opened with, on ten centimetres of board: the amplifier at 1.42 kHz, the slew limit at 7.96 kHz, Kirchhoff’s laws at 3.97 MHz and the capacitor at 4.69 MHz. Each of these is a per cent of error at the number marked and a little more just past it. Half a sample rate is not on this axis and could not be, because it is not a per cent of anything.

The difference is worth stating plainly, because it changes what an engineer does about it. Past the amplifier’s edge a per cent is available to be accepted, or two, or ten, and what it costs to move the number is a design decision with a price on it. Past half the sample rate there is no quantity to accept: the signal that was being sampled is not in the data, a signal that was not is, and the only available actions are to raise the rate or to remove the input before it arrives.

Removing it before it arrives is a filter, and a filter is an analogue circuit made of the same components as every filter elsewhere on this site — which is an essay of its own, because what it costs turns out to be a factor in the clock rate rather than a decibel or two of skirt.

A 7.0 kHz input sampled at 10 kHz arrives as 3.0 kHz. computed by solving, not by drawing. The dots are the samples. The input at 7.00 kHz is above half the 10 kHz rate, and every dot also lies on the 3.00 kHz curve drawn beside it — the two sample sequences differ by 1.6e-14, which is the arithmetic and not a small effect. Nothing is attenuated and nothing is distorted: the samples are the samples of a different signal, at full amplitude, and there is no measurement of them that could say which one was there.
Fig. 4 Seven kilohertz sampled at ten: it comes out at three. What it costs, next to the collection’s other boundaries, is that this one has no gradual failure — the input is either below half the sample rate or it is a different frequency, with nothing in between.

The hypothesis nobody meets

The sampling theorem is a conditional, and the condition is that the signal is band-limited: that it contains nothing above half the rate, not a little, not below some level. Stated that way it is a hypothesis no physical signal satisfies, and it is worth being clear about why rather than treating it as a technicality.

A signal strictly limited in frequency is unlimited in time. That is the same theorem read the other way round, and it is why every window, every switch-on and every finite measurement has a spectrum that goes on forever. A tone burst that lasts a millisecond has energy at every frequency there is; so does a step, so does anything that ever started. The band-limited signal is therefore an idealisation of exactly the kind this collection exists to put a number on — and the number is not a frequency but a level.

What the filter in front actually does is push the out-of-band content below something: below the converter’s own quantisation floor, or below the noise the source resistance already contributes, or below whatever the measurement is prepared to tolerate. That is why the design question in the next essay but one is posed as how much attenuation by which frequency, and why the answer comes out as a clock rate rather than as a yes or no. There is no rate at which aliasing stops; there is a rate at which what folds back is smaller than what is already there.

This is the same shape as every other boundary here, one level up. A lumped model is not true below some frequency and false above it; it is one degree out at 3.97 MHz, and whether one degree matters is a question the engineer answers. A band-limited signal is not a thing that exists; a signal whose out-of-band energy sits 80 dB down is, and the eighty decibels is the whole of the specification.

Folding on purpose

There is one case where the identity is used rather than avoided, and it is worth an aside because it demonstrates that nothing in the mathematics prefers the low-frequency member of the family.

A signal occupying 70–70.1 MHz — an intermediate frequency in a radio, say — has a bandwidth of 100 kHz and a centre nowhere near direct current. Sampling it at 140.2 MHz is possible and wasteful. Sampling it at, say, 400 kHz places it in an image band, and the samples that come out are the samples of a 100 kHz-wide signal somewhere between direct current and 200 kHz — exactly, at full amplitude, with the whole modulation intact. Every argument above applies unchanged; the only thing that has changed is which member of the family the designer wanted.

Two costs come with it and both are on this site’s own axes. The filter in front is now a bandpass filter, and it must reject every other image band rather than everything above a corner, which is a harder problem with the same trade in it. And the converter’s aperture jitter is now being asked about a 70 MHz input rather than a 200 kHz one — so the ceiling drawn above, which falls at twenty decibels a decade, has fallen by fifty decibels relative to the baseband case. A technique that costs nothing in sample rate costs eight bits in the aperture, which is the trade rather than a free lunch, and it is invisible unless the boundary in time is drawn beside the boundary in frequency.

The one symptom, which is not a measurement

There is a partial defence, and it is worth knowing because it is often mistaken for a full one. An input that is not a pure tone — a real signal, with a spectrum — arrives folded, and the folded copy generally overlaps what was already there. When the overlap is severe the result usually looks wrong: a tone that moves the wrong way as the source is tuned up, a rumble that appears from nowhere, a periodic beat at the difference between the input and a multiple of the rate.

That is a genuine symptom and it is not a measurement. It depends on there being something already in the band to conflict with. A quiet band and a single strong tone above half the rate produce a clean, plausible, entirely fictitious result, and the site’s own habit applies: an assertion that has never rejected anything proves nothing, and “it looked wrong when it was wrong” is not a test.

The measurement that does exist is the one in the next essay. Reconstruct the continuous signal from the samples — properly, with the sum the sampling theorem prescribes — and compare it with two things: the input, and the image. Below half the rate those are one curve. Above it they are two, the reconstruction matches the image to the accuracy of its own truncation, and it is wrong about the input by twice the amplitude.

A 11.0 kHz input sampled at 10 kHz arrives as 1.0 kHz. computed by solving, not by drawing. The dots are the samples. The input at 11.00 kHz is above half the 10 kHz rate, and every dot also lies on the 1.00 kHz curve drawn beside it — the two sample sequences differ by 2.9e-14, which is the arithmetic and not a small effect. Nothing is attenuated and nothing is distorted: the samples are the samples of a different signal, at full amplitude, and there is no measurement of them that could say which one was there.
Fig. 5 Eleven kilohertz comes out at one. The one symptom, which is not a measurement, is that the output is a perfectly good sinusoid of the wrong frequency: nothing downstream can tell it from a genuine kilohertz input, and no amount of care afterwards recovers which it was.

Where the number comes from, and where it does not

Half the sample rate is decided by the clock and by nothing else — not by the converter’s bandwidth, not by its resolution, not by the impedance in front of it. That independence is unusual on this site, where almost every boundary is a combination of at least two component values, and it is worth saying which of the converter’s other limits are not this one.

A converter’s input bandwidth is a property of its front end, and it is normally several times half the sample rate, precisely so that the sampling instant catches the signal rather than a smoothed version of it. A wide input bandwidth is therefore not protection against aliasing; it is the opposite, and a part specified with a 500 MHz input bandwidth and a 100 MHz clock will faithfully digitise everything up to 500 MHz, folded.

Its aperture jitter is a boundary in time, and it produces an error proportional to the slope of the signal — so it gets worse with input frequency and has nothing to do with the rate at all.

A 13.0 kHz input sampled at 10 kHz arrives as 3.0 kHz. computed by solving, not by drawing. The dots are the samples. The input at 13.00 kHz is above half the 10 kHz rate, and every dot also lies on the 3.00 kHz curve drawn beside it — the two sample sequences differ by 1.3e-14, which is the arithmetic and not a small effect. Nothing is attenuated and nothing is distorted: the samples are the samples of a different signal, at full amplitude, and there is no measurement of them that could say which one was there.
Fig. 6 Thirteen kilohertz comes out at three — the same three that seven kilohertz gave. Where the number comes from is the distance to the nearest multiple of the sample rate, and where it does not come from is anything about the converter: the same folding happens in a sampled control loop, a lock-in and a stroboscope, and none of them has bits in it.

And its quantisation floor is a boundary from below, in amplitude, which the next field of this site measures against the Johnson noise of the source resistance and finds crossing at a bit count that depends on the circuit in front rather than on the converter.

Four boundaries, four independent variables — a rate, a duration, an amplitude and a frequency — and only one of them has no gradient. That is the one this essay is about, and it is the reason the field opens here rather than with the arithmetic of levels: a converter’s resolution is a number to be traded against, and its sample rate is a number that is either respected or is not.

The same boundary in a circuit nobody calls a converter

The identity above is usually filed under digital signal processing, and the filters field has a circuit that runs into it while having no digital anything in it at all — no clock recovering data, no code, no arithmetic, and an output that is a continuous voltage.

A resistor made of a clock is the arrangement: a capacitor shuttled between two nodes at a megahertz behaves as a megohm, and a tenth of a picofarad shuttled at ten kilohertz behaves as a gigohm, which is how a filter with a one-hertz corner fits on a chip. What the equivalence costs is three conditions binding three different quantities, and one of the three is this essay’s — a signal below half the clock.

The filter that samples is where that condition is measured rather than stated, and the number it produces is the sharpest illustration of this essay’s claim available anywhere in the collection. An input at 992 kilohertz arrives at the output at 7.8 kilohertz with the passband’s own gain, where the continuous model that describes the circuit everywhere else says it is 56 decibels down. Fifty-six decibels of stopband, correctly designed, correctly built, correctly measured — and completely absent, because the object being described is not a continuous system and nothing below half the clock reveals that.

Two things about that are worth carrying back here. The first is that the failure is not a degradation of the filter: the passband gain is exactly right, which is what makes the alias indistinguishable from a signal that belongs there. That is this essay’s identity, arriving in a circuit whose designer was thinking about corner frequencies rather than about sample rates. The second is that the repair is the same one a converter needs and is the same essay — what the filter in front costs prices it, and finds the choice is not a decibel or two of skirt but a factor in the clock: Bessel demands 3.53 times Nyquist, Butterworth 2.08, Chebyshev 1.53, with an elliptic design refused outright because a stopband floor is not a slope and no sample rate reaches past a floor.

So the boundary this essay draws is not a property of converters. It is a property of anything that samples, and a circuit can sample without anybody having decided that it does.

What the figures do not show

Two absences are deliberate.

Nothing here samples a signal with a spectrum. Every figure in this essay samples a sinusoid, because a sinusoid is the case where the claim can be made exactly and checked exactly. A real input’s folded copy lands on top of whatever else is in the band, and the interesting quantity then is how much of it there is — which is set by the filter in front, and is the next essay but one.

Nothing here reconstructs at a rate other than the one it sampled at. Rate conversion, interpolation and decimation are all downstream of the same identity and none of them changes it; they are also the point at which this subject stops being about a circuit and starts being about an algorithm, which is a boundary this collection has stated and does not cross.

What is here is one identity, generated twice and differenced, at the arithmetic’s floor: the sharpest edge in this collection, and the only one that is invisible from the data on either side of it.

Part 1 on sample rate

One argument about Sample rate, and one of 2 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.

What this makes readable

Essays that name this one as a prerequisite.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

AliasingModel rangeNyquist rateSample rate