Filters, measured not tabulated

Three families, one corner

Butterworth is flat, Chebyshev is steep, Bessel has good delay. None of those is a number, so the table they appear in cannot answer the question anybody has. Here each family's poles are computed from its definition, built as an actual network, and then measured — starting with the step every comparison skips.

The comparison between the filter families is one of the most reproduced tables in engineering. Butterworth: maximally flat. Chebyshev: steeper, with ripple. Bessel: best transient behaviour. Every word of it is true, and none of it is a quantity, so the table cannot answer the only question a designer actually arrives with — how much steeper, and what does it cost?

Three filter families at order 5, all with the same half-power pointAt three times the corner the Chebyshev is -64.0 dB down, the Butterworth -47.7 dB and the Bessel -28.3 dB. The inset is the passband at forty times the vertical magnification, which is the only place the Chebyshev's half-decibel of ripple is visible at all.-90-60-3001001k10kfrequency (hertz)gain (decibels)ButterworthChebyshevBesselhalf power1.00 kHzthe passband, magnified-1-0.500000.2000.4000.6000.8001solved, then checked — three networks, 133 frequencies eachall normalised to a measured −3 dB at 1.00 kHz
Fig. 1 The three families at order five, all with the same measured half-power frequency. At three times the corner the Chebyshev is 64 dB down, the Butterworth 48 dB and the Bessel 28 dB. The inset is the passband at forty times the vertical magnification, which is the only place the Chebyshev’s half decibel of ripple is visible at all. The slider is the order.

Computed, then built, then measured

Each family is a rule for placing poles, and the rules are short enough to state.

Butterworth puts them evenly spaced on a semicircle in the left half-plane. That is the whole definition, and the “maximally flat” property — that as many derivatives of the magnitude as possible vanish at zero frequency — is a consequence rather than a construction.

Chebyshev puts them on the same angles but on an ellipse, whose eccentricity is set by the allowed ripple. Squash the circle toward the imaginary axis and the poles move closer to it, which makes the response peakier and steeper at once.

Bessel is the odd one: its poles are the roots of a reverse Bessel polynomial, chosen so that the phase is as close to linear as possible rather than the magnitude being as flat as possible. Its poles do not lie on any simple curve, and they are found here by rooting the polynomial numerically.

Those poles are then realised as a network — a cascade of series-resonant sections separated by ideal followers, with real element values — and every number in this field is measured on that network rather than evaluated from the polynomial. The two are required to agree to about a part in 1014 at every frequency, which is the check that the synthesis synthesised what it was asked for. A mistake in the pole placement, in the pairing of conjugates or in the conversion from a pole pair to a resistance and an inductance would break that agreement and nothing else would.

The step every comparison skips

Before any of those responses can be compared, they have to be drawn at the same place, and that is where most published comparisons go wrong.

Each family is naturally normalised to a different thing. Butterworth’s definition puts its poles on a unit circle, which happens to place its half-power point at unity. Chebyshev’s is normalised to the edge of its ripple band, which is not its half-power point — at order 5 with half a decibel of ripple, the half-power point is six per cent higher. Bessel’s classical normalisation is for unit delay, which puts its half-power point somewhere else again.

Comparing the three as they come out of their definitions therefore compares three curves drawn at three different frequencies, and the result flatters or damns each family for reasons that have nothing to do with its behaviour. It is a large part of why Bessel filters have a reputation for being much shallower than they are.

So here every family is rescaled until its own measured half-power point lands at the same frequency, and the measurement is a bisection on the solved response rather than a factor from a table. The assertion in the figure is that all three land within three parts in a thousand of the target, which is the resolution of the bisection rather than a tolerance on the design.

The numbers, at order five

With that done, the comparison means something:

at 3×f_c passband deviation delay variation
Butterworth −47.7 dB 0.44 dB 48.0%
Chebyshev −64.0 dB 0.50 dB 49.0%
Bessel −28.3 dB 1.88 dB 0.06%

Three things in that table are worth stopping on.

The Chebyshev’s advantage is 16 dB. That is a factor of six in amplitude, at three times the corner, for the same number of components — a substantial and entirely real benefit, and it is now a number rather than the word “steeper”.

The Bessel’s disadvantage is larger than its advantage looks. It is 19 dB behind the Butterworth in the stopband, which is a factor of nine. Anyone choosing it for its transient behaviour is paying a great deal for it, and should know how much.

The passband column does not say what the adjectives suggest. The Butterworth’s deviation is 0.44 dB and the Chebyshev’s is 0.50 dB — essentially the same. “Maximally flat” is a statement about derivatives at zero frequency, not about how flat the passband is overall, and by the second measure the rippled filter is no worse. The Bessel’s 1.88 dB, meanwhile, is the largest of the three: its magnitude sags gently across the whole band, which is the price of a linear phase and is not described by the word “flat” at all.

What changes with the order

The slider moves the order from two to eight, and the trade moves with it in a way worth reading off rather than deriving.

The far stopband slope is 20 n decibels per decade for every family — a hundred at order five, a hundred and forty at order seven — because all three are all-pole filters of the same order and far enough out nothing else matters. What differs is how quickly each reaches that slope, and that is what the measurement at three times the corner captures.

The gap between families widens with order. At order three the Chebyshev leads the Bessel by 14 dB; at order five by 36; at order seven by 62. The choice of family therefore matters more, not less, as the filter gets more complicated — which is the opposite of the usual intuition that a high-order filter approaches an ideal brick wall regardless of how it is made.

And the Bessel’s delay variation falls with order, from 2.3% at order three to 0.06% at five and below 0.01% at seven, while the other two stay in the region of fifty per cent. That is the family doing exactly what it was designed for, and it is the only column in the table that improves with complexity.

What each family costs, at order 5Measured on the solved networks. The Chebyshev is 36 dB further down at three times the corner than the Bessel, and pays for it in delay: its group delay varies 49.0% across the passband against the Bessel's 0.06%.passband deviationdecibels, peak to trough below 0.8 f_cButterworth0.443 dBChebyshev0.500 dBBessel1.882 dBattenuation at three times the cornerdecibels downButterworth47.7 dBChebyshev64.0 dBBessel28.3 dBgroup-delay variation across the passbandper cent, slowest against fastestButterworth48.0%Chebyshev49.0%Bessel0.1%solved, then checked — nine measurements, three networksevery number here moves with the order
Fig. 2 The same three quantities as bars rather than a table, so the shape of the trade is visible at a glance. Nothing here is quoted; every bar is a measurement on the solved network, and every one of them moves when the order does.

Where the poles sit, and why the shapes follow

The three responses in the figure look different because the three sets of poles are in different places, and the correspondence is direct enough to be worth spelling out.

Butterworth’s poles are evenly spaced on a semicircle. All of them are the same distance from the origin, so all have the same natural frequency, and the ones nearest the imaginary axis — the pair at the ends of the arc — have a quality factor that rises with order. At order five that pair has a quality factor of about 1.6; at order nine, about 2.9. That pair is what produces both the sharpness of the transition and the peak in group delay.

Chebyshev’s poles are on an ellipse squashed toward the imaginary axis. Every pole is therefore closer to the axis than the corresponding Butterworth pole, the extreme pair much more so — at order five with half a decibel of ripple its quality factor is about 8.8, five times the Butterworth’s. A pole with a quality factor of nearly nine is a sharp resonance, and the passband ripple is that resonance showing through.

Bessel’s poles are spread along a curve that keeps them away from the axis. No pair has a high quality factor at any order, which is why nothing about the response is sharp: not the transition, not the delay, not the step.

So the three families are one statement made three ways. Everything sharp in a filter is a pole near the imaginary axis, and everything a sharp filter costs is what such a pole does in time. The comparison in this essay is a measurement of that single trade at three points along it.

What the realisation assumes

The networks measured here separate their sections with ideal followers, and the assumption deserves naming because it is a choice with consequences rather than a simplification.

With followers between them, each second-order section’s response is its own, and the cascade’s poles are exactly the union of the sections’ poles. Without them — in a passive ladder, which is the usual realisation at radio frequencies — each section loads the one before it, every pole moves, and the element values have to come from a different synthesis that accounts for the loading.

Both realise the same response, so nothing in the measurements changes. What changes is sensitivity: how much the response moves when a component is not quite its nominal value. A cascade of isolated sections has each pole determined by two or three components, so a one per cent error in one component moves one pole by about one per cent. A doubly-terminated passive ladder has a remarkable property — at frequencies where it transfers maximum power, the response is stationary with respect to every element value, so first-order errors cancel — which is why passive ladders are used where components are the limiting factor.

That is a genuine advantage of a realisation the figures here do not use, and it is worth recording as a limitation of the comparison rather than left for a reader to discover.

What is deliberately missing

There is a fourth family in every textbook version of this comparison, and it is not here.

The elliptic filter places zeros on the imaginary axis as well as poles, giving it a stopband that ripples rather than descending monotonically and a skirt steeper than a Chebyshev of the same order. It is the right answer whenever the requirement is a hard transition and the stopband only has to be below something rather than getting steadily further below it.

It is left out for a reason worth stating rather than hidden. Placing its zeros requires Jacobi elliptic functions, and realising them requires a ladder with resonant branches rather than the cascade of simple sections used here — a different and considerably longer piece of machinery. It sits further along the same axis as Chebyshev, steeper still and worse still in delay, so the three families here bracket the trade without it.

That is an honest gap rather than an oversight, and it is recorded as one in the library that builds these filters. A collection that quietly drops the inconvenient case and does not say so is doing something different from one that says what it left out.

The order that is not an integer

One practical note about the slider, since it moves through integers and a specification rarely lands on one.

The order of an all-pole filter is the number of poles, so it is a whole number by construction. A requirement that needs 4.3 poles has to be met with five, and the resulting filter exceeds the specification — usually by a comfortable margin, since each additional pole is worth twenty decibels per decade in the stopband.

That granularity is coarser than it looks. Between order four and order five, the attenuation at three times the corner improves by about ten decibels for a Butterworth and fifteen for a Chebyshev. So the choice is rarely between meeting the specification and missing it; it is between one design that misses and another that overshoots substantially, and the overshoot is paid for in components, cost, and — as this field keeps returning to — delay.

Which is one more argument for measuring the third column. If order five overshoots the magnitude requirement anyway, the freedom that buys can be spent on a gentler family: a Bessel of order six may meet the same stopband requirement as a Chebyshev of order four while having a hundredth of the delay variation. That trade is available only to somebody who has both numbers, and it is invisible in a table of adjectives.

The two routes, and what they caught

Every filter here exists twice: as a set of poles computed from a definition, and as a netlist of resistors, inductors, capacitors and followers. The figure asserts that the response of the second matches the response of the first to about a part in 1014 across four decades.

That agreement is not automatic, and the ways it can fail are instructive. Pairing the wrong two poles as conjugates produces sections with complex element values, which the network refuses. Getting the quality factor of a section wrong by a factor of two — easy, since the definition involves both a factor of two and a square root — produces a response that is still a low-pass of the right order with roughly the right corner, and is a different filter. Omitting the normalisation step produces three perfectly correct filters drawn at three different frequencies, which is the published state of the art.

Only the first of those three is caught by the network refusing. The other two are caught by comparing routes, and they are exactly the mistakes that would otherwise survive into a figure that looked entirely convincing.

Group delay across the passband, at order 5The Bessel filter's delay varies 0.1% below 0.8 of the corner; the Chebyshev's peaks near the band edge and is several times its low-frequency value. Every family here has the same half-power frequency, so this is a difference in behaviour rather than in scaling.00.5011.521001kfrequency (hertz)group delay (milliseconds)ButterworthChebyshevBesselthe corner, 1.00 kHzsolved, then checked — −dφ/dω on the unwrapped phaseflat magnitude is not flat delay
Fig. 3 The column of the table that the classical comparison omits, drawn out. Group delay against frequency for the same three networks: the Bessel is a horizontal line and the other two rise steeply toward the band edge. What that does to a signal is the subject of a later essay in this field.
The same step through all three, at order 5Overshoot measured off each curve: Butterworth 12.8%, Chebyshev 12.4%, Bessel 0.8%. The family with the flat delay barely overshoots; the other two, whose delay varies by tens of per cent, overshoot by more than ten and are hard to tell apart — magnitude flatness does not decide it.00.500101234time (milliseconds)output, for a 1 V step inButterworth 12.8%Chebyshev 12.4%Bessel 0.8%solved, then checked — residues, checked by integrationovershoot follows the delay, not the magnitude
Fig. 4 What the three do to a step, which is the measurement the magnitude plots cannot supply. The Chebyshev and the Butterworth are almost indistinguishable here despite being the furthest apart in the stopband; the Bessel, which is the worst of the three by every magnitude measure, is the only one that arrives cleanly.

Choosing, in the terms the measurements support

The point of measuring is to be able to state the choice without adjectives, and it comes out roughly as follows.

If the requirement is a specified attenuation at a specified frequency and nothing else matters, the Chebyshev wins by a margin that grows with order, and the ripple it costs is invisible on any plot that is not magnified forty times. If the signal has edges in it — a pulse, a data stream, a step — and the shape matters, the Bessel is the only one of the three whose delay is flat, and everything else about it is worse. The Butterworth is the compromise, and its real distinction is not its flatness but that it has no bad property at all: no ripple, no sag, no extreme delay variation, nothing that will surprise anybody.

Which is, in the end, why it is the default. Not because maximal flatness is usually the requirement, but because a filter with no distinguishing weakness is the right choice when the requirement has not been stated precisely — and most of the time it has not.