Filters, measured not tabulated

Flat magnitude, unflat delay

A filter that passes every frequency in its band at the right amplitude and the wrong time has not passed the signal. Group delay is the measurement that says so, it is absent from the classical comparison, and it varies by fifty per cent across the passband of the two families everybody uses.

Assumes: Three families, one corner · One step, computed twice

A filter is normally judged by what it does to amplitudes. That is one of two things it does, and for anything with edges in it the other one matters more: a filter also delays, and it does not delay every frequency by the same amount.

Group delay across the passband, at order 5The Bessel filter's delay varies 0.1% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.00.5011.521001kfrequency (hertz)group delay (milliseconds)ButterworthChebyshevBesselthe corner, 1.00 kHzsolved, then checked — −dφ/dω on the unwrapped phaseflat magnitude is not flat delay
Fig. 1 Group delay against frequency for the three families at order five, all with the same half-power point. The Bessel is a horizontal line — 0.06% variation across the passband. The other two rise steeply toward the band edge, by about half again. The slider is the order.

What group delay measures

A filter shifts the phase of every frequency it passes. If it shifted them all by the same angle the waveform would still be wrecked, because a fixed angle is a different fraction of a cycle at every frequency. What leaves a waveform intact is a fixed time, which means a phase shift proportional to frequency — a straight line through the origin on a plot of phase against frequency.

Group delay is the slope of that line, computed locally:

τ(ω) = −dφ/dω

A perfectly linear phase gives a constant group delay, and a constant group delay means every component of a signal is held up by the same amount and reassembled in the right order at the far end. The signal comes out late and otherwise unaltered. Where the group delay is not constant, components arrive at different times, and what comes out is not the signal shifted — it is a different shape.

Why it is absent from the usual comparison

The reason group delay is missing from the classical filter table is worth stating, because it is not carelessness.

The classical treatment grew up around problems where it genuinely does not matter. Separating one radio channel from its neighbours, removing mains hum from a measurement, band-limiting before sampling a slowly varying quantity: in all of those the signal of interest is narrow-band, and over a narrow band any smooth group delay is approximately constant. What the filter does to the shape of a wide-band transient is not asked about, because there is no wide-band transient.

The moment the signal has edges — a data stream, a pulse, a step, anything switched — the question becomes the central one, and the table has nothing to say about it.

The measurement, and the step that makes it work

Group delay is a derivative of phase, and phase comes out of an arctangent, which means it comes out wrapped into a range of 360 degrees. Differentiating a wrapped phase produces an enormous spike at every wrap: a curve with a set of vertical spikes in it that look exactly like resonances and are entirely artefacts of the arithmetic.

So the phase is unwrapped first — walked along the sweep, adding or subtracting full turns to keep it continuous — and only then differentiated. The computation here uses a central difference over a frequency interval of a part in a thousand, with the unwrapping done locally so that a wrap between the two sample points cannot survive into the answer.

That is a small detail with a large consequence, and it is the kind of thing worth naming because the failure it prevents produces a plausible-looking figure. A group-delay plot with spikes in it would be read as a filter with problems, and the problems would be in the plotting.

The check that the whole computation is sound is a physical one: the group delay must be positive everywhere. A negative group delay in a passive filter would be a network responding before its input arrived. The figure asserts it at every point of every curve, and it is the one assertion here that could catch an unwrapping failure, a sign error and a mislabelled axis at once.

What the difference does to a step

A step contains every frequency. A filter whose group delay varies across the band cannot reassemble it, and what comes out overshoots and rings.

The same step through all three, at order 5. Overshoot measured off each curve: Butterworth 12.8%, Chebyshev 12.4%, Bessel 0.8%. The family with the flat delay barely overshoots; the other two, whose delay varies by tens of per cent, overshoot by more than ten and are hard to tell apart — magnitude flatness does not decide it.
Fig. 2 The same step through all three, at the same order and the same corner. The Bessel arrives late, cleanly, and stops. The other two overshoot by more than ten per cent and ring for several cycles. The connection to the previous figure is the whole point: nothing in the magnitude plots distinguishes these three curves.
Group delay across the passband, at order 4. The Bessel filter's delay varies 0.4% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.
Fig. 3 The same three families at order four. The Bessel’s delay varies 0.4% below 0.8 of the corner, against 0.1% at order five, and the Chebyshev’s still peaks near the band edge at several times its low-frequency value. All three share a half-power frequency, so what is drawn is a difference in behaviour rather than a difference in scaling.

The measured overshoots at order five are Butterworth 12.8%, Chebyshev 12.4%, Bessel 0.8%.

That grouping is worth reading carefully, because it is not the grouping the magnitude plots suggest. The Butterworth and the Chebyshev — the flattest and the rippliest — behave almost identically on a step, differing by less than half a percentage point. The Bessel, whose magnitude is the least flat of the three by the measurements in the previous essay, overshoots sixteen times less.

So passband flatness does not predict transient behaviour, at all. The quantity that does is the delay variation, and the two families with fifty per cent of it are indistinguishable while the one with 0.06% is in a different category.

Phase delay is a different quantity

Two delays are defined for a filter and conflating them is a common source of confusion, so it is worth separating them once.

Phase delay is −φ/ω: how much a single sinusoid at that frequency is held up. Group delay is −dφ/dω: how much the envelope of a narrow band of frequencies around that point is held up. For a perfectly linear phase the two are equal and constant, which is why the distinction can be ignored in the ideal case and not otherwise.

The one that matters for a signal is the group delay, because a signal is a band rather than a sinusoid, and it is the envelope that carries the information. A filter can have a badly behaved phase delay and a flat group delay — an all-pass network is exactly that, with a phase that varies enormously and a delay that can be made constant — and it will pass a waveform intact.

There is also a useful sanity check hiding in the definitions. At zero frequency the two must agree, because a phase proportional to frequency and its own derivative coincide at the origin. Every curve in the figure above therefore starts at the same value as the phase delay of the same filter at low frequency, and a computation that got the differentiation wrong would break that agreement at the left-hand edge before it broke anything else.

Where the numbers land in practice

The abstraction is easier to hold with an application attached, and two are worth naming because they sit at opposite ends of the trade.

An anti-aliasing filter before a sampler cares about the stopband and essentially nothing else. Anything above half the sampling rate folds back into the band irreversibly, so the requirement is an attenuation at a frequency, the signal is usually band-limited and slowly varying, and the delay across the passband is uniform enough not to matter. This is the Chebyshev’s case, and it is where the classical table’s advice is exactly right.

A pulse-shaping filter in a data link cares about almost nothing else but the delay. The signal is a sequence of edges; the whole point is that each symbol should be recoverable at its own instant; and a filter that delays the high-frequency content of an edge differently from the low smears one symbol into the next. That smearing has a name — intersymbol interference — and it is measured in exactly the terms this essay uses. Here the Bessel’s twenty decibels of lost stopband are cheap and the Chebyshev is unusable at any order.

The interesting cases are the ones in between, and the reason a measured comparison is worth having rather than an adjective is that they cannot be settled by a rule. A filter in an instrumentation chain, or in an audio crossover, or ahead of a control loop, has requirements on both axes, and the only way to know whether a given order and family meets them is to have both numbers.

The claim this figure does not make

There is a tempting next sentence — that the overshoot ordering follows the delay-variation ordering exactly — and it is not true, so the figure does not assert it.

At order five the Chebyshev’s delay varies slightly more than the Butterworth’s (49.0% against 48.0%) and it overshoots slightly less (12.4% against 12.8%). At order three the two swap again. The relation between the two quantities is monotone across the large gap and noisy across the small one, which is what should be expected: the overshoot depends on the whole shape of the delay curve and the summary statistic is one number taken from it.

What the figure does assert, at every order it can be dragged to, is that the flat-delay family overshoots at least three times less than the maximally flat one. That is the claim the measurements support, and it is stated in the code rather than in a caption so that a future change to the synthesis cannot quietly break it.

Stopping at the claim the data supports, rather than the tidier one next to it, is a small discipline that matters here more than usual — because the tidier claim is exactly what a reader would expect and would therefore accept without checking.

The unwrapping, and what it costs to get wrong

It is worth spending a paragraph on the mechanics, because this is the one measurement in the collection whose commonest failure produces a figure that looks like a discovery.

Phase comes out of a two-argument arctangent, which returns a value in a range of 360 degrees. A sweep across a filter’s response passes through several multiples of that range — a fifth-order low-pass accumulates 450 degrees of phase — so the raw phase curve has four discontinuities in it, each a jump of exactly one turn. Differentiating that raw curve gives an enormous positive or negative spike at each jump.

Those spikes are narrow, tall, and sit at frequencies determined by the filter’s own poles, which makes them look exactly like resonances. A reader shown such a plot would reasonably conclude the filter had four sharp features in its delay. It has none; the arctangent has four.

The remedy is to walk the sweep adding or subtracting whole turns to keep the phase continuous before differentiating, which is what “unwrapping” means. The version used here goes further and unwraps locally, between the two points of each central difference, so that a wrap falling between two adjacent samples cannot survive into the derivative even if the global unwrapping were to slip.

And then the check, which is what makes the whole thing trustworthy rather than merely careful: the group delay must be positive everywhere. A negative group delay in a passive network is a network responding to an input before it arrives. The figure asserts it at every point of every curve, and that single assertion would catch a failed unwrapping, a sign error in the difference, an inverted frequency axis and a transposed pair of arguments — four distinct mistakes, all of which produce a plausible-looking curve.

Where the delay goes

A last observation that makes the figures easier to read.

Group delay is largest where poles are closest to the imaginary axis, because a pole close to the axis is a sharp feature in frequency and a sharp feature in frequency is a long one in time. That is the same statement as the bandwidth–settling-time relation in the transients field, applied to each pole pair individually.

It explains the shape of every curve in the top figure. Butterworth and Chebyshev both have a pole pair near the band edge with a high quality factor — that is what makes their skirts steep — and that pair contributes a peak in group delay just below the corner. Bessel’s poles are spread out and none of them is especially close to the axis, which is why its skirt is gentle and its delay is flat. The steep skirt and the delay peak are the same pole.

Which means the trade is structural rather than a matter of design skill. No arrangement of poles gives both a sharp transition and a flat delay, because the sharp transition is a pole near the axis and a pole near the axis is a delay peak. A designer can choose where on that curve to sit; nobody can leave it.

Group delay across the passband, at order 2. The Bessel filter's delay varies 11.2% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.
Fig. 4 Order two, the lowest the comparison is drawn at. The Bessel’s delay varies 11.2% across the passband — its flatness is a property of high order, not of the family — and at this order the three families are much closer together than the essay’s headline suggests.

Why the ear is the exception

One qualification belongs here because it is the source of a long argument and the measurements above do not settle it by themselves.

For most signals, delay variation is straightforwardly damaging: an edge arrives smeared and a data symbol runs into its neighbour. For audio, the case is weaker than the numbers suggest, because hearing is substantially insensitive to absolute phase. A tone burst passed through a filter with fifty per cent of delay variation is measurably different at the output and is very often not audibly different, and the threshold at which it becomes audible has been measured and is around one to two milliseconds of variation across the band — far more than most crossovers produce.

That does not make the measurement irrelevant, and the reason is worth stating precisely. The delay variation still exists, it is still what produces the overshoot on a step, and in a loudspeaker crossover the overshoot is a real acoustic event rather than a curve. What it does mean is that the criterion is different in audio: the question is not whether the delay is flat but whether its variation is below a threshold set by perception rather than by arithmetic.

This collection has nothing to add about perception, and the honest thing is to say so and hand the number over. What it can supply is the variation in milliseconds rather than per cent, which is the form a perceptual threshold is stated in — and which any reader can take from the caption strip of the figure at the top of this essay.

Group delay across the passband, at order 3. The Bessel filter's delay varies 2.3% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.
Fig. 5 Order three: the Bessel is down to 2.3%. Between orders two and four the variation falls 11.2%, 2.3%, 0.4% — a factor of about five per order, which is the sense in which a Bessel filter is designed for flat delay rather than merely having it.
Group delay across the passband, at order 8. The Bessel filter's delay varies 0.0% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.
Fig. 6 And order eight, the end of the slider, where the Bessel’s passband delay variation rounds to 0.0% and the Chebyshev’s peak has grown. The two families move in opposite directions with order, which is why “a steeper filter” is not a single specification: it is steeper in magnitude and worse in delay unless the family was chosen for the second.

The one that gets it both ways, and what it costs

Since the trade is structural, the only way out is to stop trying to do it with poles alone.

A filter’s delay can be flattened after the fact by an all-pass section: a network with unity magnitude at every frequency and a phase that varies, added specifically to make the total phase linear. It works, it is what a well-engineered signal chain does, and it is entirely absent from the classical table.

The cost is components — an all-pass equaliser for a fifth-order filter is typically two or three more sections, so more than the filter itself — and, unavoidably, more delay. Equalising a group delay means adding delay at the frequencies that were fastest until they match the slowest, so the result is flat at a value at least as large as the original maximum. Signal shape is bought with latency, at a fixed rate, and no arrangement improves the exchange.

That is the last of the three currencies in this field. Attenuation is bought with poles, at twenty decibels per decade each. Steepness is bought with delay variation, by moving poles toward the imaginary axis. And flat delay is bought back with latency, by adding all-pass sections. Every one of those is a rate rather than an adjective, and the figures in this field exist to put a number on each.

Flat delay, bought with more delay is where the third rate is measured rather than described, and two of its numbers change how the trade reads. The all-pass section’s magnitude is one to within 9×10169\times10^{-16} over six decades on a solved network — which is what entitles anyone to treat delay and magnitude as separable at all — and the exchange on a fifth-order Chebyshev is 891 microseconds of delay variation coming down to 498, with everything leaving 722 microseconds later than it did. So the flattening is a factor of 1.8 rather than a correction, and the latency paid for it is larger than the variation removed.

What a steep skirt costs is where the second rate is measured, and its result is the one that makes this essay’s fifty per cent look mild: across the three families at one order, the steepest distorts delay eight hundred times more than the gentlest. That is a ratio between families rather than across a passband, and it is the number a designer choosing a family is actually spending.

The first rate is the one nothing changes, and it is worth restating for that reason. Twenty decibels per decade per pole is a property of a rational function and not of a circuit, so no realisation, no family and no arrangement of components moves it — which is why every argument in this field is about what the other two currencies buy.

The one family that appears to escape it does not. The zeros that buy an order reaches forty decibels at twice the corner with two poles where the steepest all-pole family needs five, by putting zeros in the stopband — which changes how quickly the slope is reached rather than the slope itself, and charges for it with a stopband that stops falling at −38.7 dB.

And the delay this essay measures has a provenance worth naming. The phase the magnitude already knows shows that for a minimum-phase network the phase — and therefore the group delay, which is its derivative — is fixed everywhere by the magnitude. So a filter’s delay variation is not an independent property to be traded against its magnitude response: it is implied by it, and the only way to change one without the other is to leave the class, which is what an all-pass section does.

What the delay costs, and what is done about it

A group delay that varies across the passband is the defect three later essays in this field are about. Flat delay, bought with more delay is the repair, and it is expensive: three all-pass sections remove 600 µs of variation and add 2,043 µs of delay. Three families, one corner is where the family that needs least repair is chosen, and What a steep skirt costs is the decision that creates the problem in the first place. The staircase on the way out is the delay a converter adds that no filter design contains — exactly half a sample, at every rate. And One step, computed twice is the machinery that makes the time-domain half of all of it computable.

Part 3 on group delay

One argument about Group delay, and one of 2 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that name this one as a prerequisite.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Group delayOvershootPhase linearityRingingUnwrapping