Flat magnitude, unflat delay
A filter is normally judged by what it does to amplitudes. That is one of two things it does, and for anything with edges in it the other one matters more: a filter also delays, and it does not delay every frequency by the same amount.
What group delay measures
A filter shifts the phase of every frequency it passes. If it shifted them all by the same angle the waveform would still be wrecked, because a fixed angle is a different fraction of a cycle at every frequency. What leaves a waveform intact is a fixed time, which means a phase shift proportional to frequency — a straight line through the origin on a plot of phase against frequency.
Group delay is the slope of that line, computed locally:
τ(ω) = −dφ/dω
A perfectly linear phase gives a constant group delay, and a constant group delay means every component of a signal is held up by the same amount and reassembled in the right order at the far end. The signal comes out late and otherwise unaltered. Where the group delay is not constant, components arrive at different times, and what comes out is not the signal shifted — it is a different shape.
Why it is absent from the usual comparison
The reason group delay is missing from the classical filter table is worth stating, because it is not carelessness.
The classical treatment grew up around problems where it genuinely does not matter. Separating one radio channel from its neighbours, removing mains hum from a measurement, band-limiting before sampling a slowly varying quantity: in all of those the signal of interest is narrow-band, and over a narrow band any smooth group delay is approximately constant. What the filter does to the shape of a wide-band transient is not asked about, because there is no wide-band transient.
The moment the signal has edges — a data stream, a pulse, a step, anything switched — the question becomes the central one, and the table has nothing to say about it.
The measurement, and the step that makes it work
Group delay is a derivative of phase, and phase comes out of an arctangent, which means it comes out wrapped into a range of 360 degrees. Differentiating a wrapped phase produces an enormous spike at every wrap: a curve with a set of vertical spikes in it that look exactly like resonances and are entirely artefacts of the arithmetic.
So the phase is unwrapped first — walked along the sweep, adding or subtracting full turns to keep it continuous — and only then differentiated. The computation here uses a central difference over a frequency interval of a part in a thousand, with the unwrapping done locally so that a wrap between the two sample points cannot survive into the answer.
That is a small detail with a large consequence, and it is the kind of thing worth naming because the failure it prevents produces a plausible-looking figure. A group-delay plot with spikes in it would be read as a filter with problems, and the problems would be in the plotting.
The check that the whole computation is sound is a physical one: the group delay must be positive everywhere. A negative group delay in a passive filter would be a network responding before its input arrived. The figure asserts it at every point of every curve, and it is the one assertion here that could catch an unwrapping failure, a sign error and a mislabelled axis at once.
What the difference does to a step
A step contains every frequency. A filter whose group delay varies across the band cannot reassemble it, and what comes out overshoots and rings.
The measured overshoots at order five are Butterworth 12.8%, Chebyshev 12.4%, Bessel 0.8%.
That grouping is worth reading carefully, because it is not the grouping the magnitude plots suggest. The Butterworth and the Chebyshev — the flattest and the rippliest — behave almost identically on a step, differing by less than half a percentage point. The Bessel, whose magnitude is the least flat of the three by the measurements in the previous essay, overshoots sixteen times less.
So passband flatness does not predict transient behaviour, at all. The quantity that does is the delay variation, and the two families with fifty per cent of it are indistinguishable while the one with 0.06% is in a different category.
Phase delay is a different quantity
Two delays are defined for a filter and conflating them is a common source of confusion, so it is worth separating them once.
Phase delay is −φ/ω: how much a single sinusoid at that frequency is held up. Group delay is −dφ/dω: how much the envelope of a narrow band of frequencies around that point is held up. For a perfectly linear phase the two are equal and constant, which is why the distinction can be ignored in the ideal case and not otherwise.
The one that matters for a signal is the group delay, because a signal is a band rather than a sinusoid, and it is the envelope that carries the information. A filter can have a badly behaved phase delay and a flat group delay — an all-pass network is exactly that, with a phase that varies enormously and a delay that can be made constant — and it will pass a waveform intact.
There is also a useful sanity check hiding in the definitions. At zero frequency the two must agree, because a phase proportional to frequency and its own derivative coincide at the origin. Every curve in the figure above therefore starts at the same value as the phase delay of the same filter at low frequency, and a computation that got the differentiation wrong would break that agreement at the left-hand edge before it broke anything else.
Where the numbers land in practice
The abstraction is easier to hold with an application attached, and two are worth naming because they sit at opposite ends of the trade.
An anti-aliasing filter before a sampler cares about the stopband and essentially nothing else. Anything above half the sampling rate folds back into the band irreversibly, so the requirement is an attenuation at a frequency, the signal is usually band-limited and slowly varying, and the delay across the passband is uniform enough not to matter. This is the Chebyshev’s case, and it is where the classical table’s advice is exactly right.
A pulse-shaping filter in a data link cares about almost nothing else but the delay. The signal is a sequence of edges; the whole point is that each symbol should be recoverable at its own instant; and a filter that delays the high-frequency content of an edge differently from the low smears one symbol into the next. That smearing has a name — intersymbol interference — and it is measured in exactly the terms this essay uses. Here the Bessel’s twenty decibels of lost stopband are cheap and the Chebyshev is unusable at any order.
The interesting cases are the ones in between, and the reason a measured comparison is worth having rather than an adjective is that they cannot be settled by a rule. A filter in an instrumentation chain, or in an audio crossover, or ahead of a control loop, has requirements on both axes, and the only way to know whether a given order and family meets them is to have both numbers.
The claim this figure does not make
There is a tempting next sentence — that the overshoot ordering follows the delay-variation ordering exactly — and it is not true, so the figure does not assert it.
At order five the Chebyshev’s delay varies slightly more than the Butterworth’s (49.0% against 48.0%) and it overshoots slightly less (12.4% against 12.8%). At order three the two swap again. The relation between the two quantities is monotone across the large gap and noisy across the small one, which is what should be expected: the overshoot depends on the whole shape of the delay curve and the summary statistic is one number taken from it.
What the figure does assert, at every order it can be dragged to, is that the flat-delay family overshoots at least three times less than the maximally flat one. That is the claim the measurements support, and it is stated in the code rather than in a caption so that a future change to the synthesis cannot quietly break it.
Stopping at the claim the data supports, rather than the tidier one next to it, is a small discipline that matters here more than usual — because the tidier claim is exactly what a reader would expect and would therefore accept without checking.
The unwrapping, and what it costs to get wrong
It is worth spending a paragraph on the mechanics, because this is the one measurement in the collection whose commonest failure produces a figure that looks like a discovery.
Phase comes out of a two-argument arctangent, which returns a value in a range of 360 degrees. A sweep across a filter’s response passes through several multiples of that range — a fifth-order low-pass accumulates 450 degrees of phase — so the raw phase curve has four discontinuities in it, each a jump of exactly one turn. Differentiating that raw curve gives an enormous positive or negative spike at each jump.
Those spikes are narrow, tall, and sit at frequencies determined by the filter’s own poles, which makes them look exactly like resonances. A reader shown such a plot would reasonably conclude the filter had four sharp features in its delay. It has none; the arctangent has four.
The remedy is to walk the sweep adding or subtracting whole turns to keep the phase continuous before differentiating, which is what “unwrapping” means. The version used here goes further and unwraps locally, between the two points of each central difference, so that a wrap falling between two adjacent samples cannot survive into the derivative even if the global unwrapping were to slip.
And then the check, which is what makes the whole thing trustworthy rather than merely careful: the group delay must be positive everywhere. A negative group delay in a passive network is a network responding to an input before it arrives. The figure asserts it at every point of every curve, and that single assertion would catch a failed unwrapping, a sign error in the difference, an inverted frequency axis and a transposed pair of arguments — four distinct mistakes, all of which produce a plausible-looking curve.
Where the delay goes
A last observation that makes the figures easier to read.
Group delay is largest where poles are closest to the imaginary axis, because a pole close to the axis is a sharp feature in frequency and a sharp feature in frequency is a long one in time. That is the same statement as the bandwidth–settling-time relation in the transients field, applied to each pole pair individually.
It explains the shape of every curve in the top figure. Butterworth and Chebyshev both have a pole pair near the band edge with a high quality factor — that is what makes their skirts steep — and that pair contributes a peak in group delay just below the corner. Bessel’s poles are spread out and none of them is especially close to the axis, which is why its skirt is gentle and its delay is flat. The steep skirt and the delay peak are the same pole.
Which means the trade is structural rather than a matter of design skill. No arrangement of poles gives both a sharp transition and a flat delay, because the sharp transition is a pole near the axis and a pole near the axis is a delay peak. A designer can choose where on that curve to sit; nobody can leave it.
Why the ear is the exception
One qualification belongs here because it is the source of a long argument and the measurements above do not settle it by themselves.
For most signals, delay variation is straightforwardly damaging: an edge arrives smeared and a data symbol runs into its neighbour. For audio, the case is weaker than the numbers suggest, because hearing is substantially insensitive to absolute phase. A tone burst passed through a filter with fifty per cent of delay variation is measurably different at the output and is very often not audibly different, and the threshold at which it becomes audible has been measured and is around one to two milliseconds of variation across the band — far more than most crossovers produce.
That does not make the measurement irrelevant, and the reason is worth stating precisely. The delay variation still exists, it is still what produces the overshoot on a step, and in a loudspeaker crossover the overshoot is a real acoustic event rather than a curve. What it does mean is that the criterion is different in audio: the question is not whether the delay is flat but whether its variation is below a threshold set by perception rather than by arithmetic.
This collection has nothing to add about perception, and the honest thing is to say so and hand the number over. What it can supply is the variation in milliseconds rather than per cent, which is the form a perceptual threshold is stated in — and which any reader can take from the caption strip of the figure at the top of this essay.
The one that gets it both ways, and what it costs
Since the trade is structural, the only way out is to stop trying to do it with poles alone.
A filter’s delay can be flattened after the fact by an all-pass section: a network with unity magnitude at every frequency and a phase that varies, added specifically to make the total phase linear. It works, it is what a well-engineered signal chain does, and it is entirely absent from the classical table.
The cost is components — an all-pass equaliser for a fifth-order filter is typically two or three more sections, so more than the filter itself — and, unavoidably, more delay. Equalising a group delay means adding delay at the frequencies that were fastest until they match the slowest, so the result is flat at a value at least as large as the original maximum. Signal shape is bought with latency, at a fixed rate, and no arrangement improves the exchange.
That is the last of the three currencies in this field. Attenuation is bought with poles, at twenty decibels per decade each. Steepness is bought with delay variation, by moving poles toward the imaginary axis. And flat delay is bought back with latency, by adding all-pass sections. Every one of those is a rate rather than an adjective, and the figures in this field exist to put a number on each.