Three families, one corner
Assumes: One solve, read four ways · Where the behaviour is written down
The comparison between the filter families is one of the most reproduced tables in engineering. Butterworth: maximally flat. Chebyshev: steeper, with ripple. Bessel: best transient behaviour. Every word of it is true, and none of it is a quantity, so the table cannot answer the only question a designer actually arrives with — how much steeper, and what does it cost?
Computed, then built, then measured
Each family is a rule for placing poles, and the rules are short enough to state.
Butterworth puts them evenly spaced on a semicircle in the left half-plane. That is the whole definition, and the “maximally flat” property — that as many derivatives of the magnitude as possible vanish at zero frequency — is a consequence rather than a construction.
Chebyshev puts them on the same angles but on an ellipse, whose eccentricity is set by the allowed ripple. Squash the circle toward the imaginary axis and the poles move closer to it, which makes the response peakier and steeper at once.
Bessel is the odd one: its poles are the roots of a reverse Bessel polynomial, chosen so that the phase is as close to linear as possible rather than the magnitude being as flat as possible. Its poles do not lie on any simple curve, and they are found here by rooting the polynomial numerically.
Those poles are then realised as a network — a cascade of series-resonant sections separated by ideal followers, with real element values — and every number in this field is measured on that network rather than evaluated from the polynomial. The two are required to agree to about a part in 1014 at every frequency, which is the check that the synthesis synthesised what it was asked for. A mistake in the pole placement, in the pairing of conjugates or in the conversion from a pole pair to a resistance and an inductance would break that agreement and nothing else would.
The step every comparison skips
Before any of those responses can be compared, they have to be drawn at the same place, and that is where most published comparisons go wrong.
Each family is naturally normalised to a different thing. Butterworth’s definition puts its poles on a unit circle, which happens to place its half-power point at unity. Chebyshev’s is normalised to the edge of its ripple band, which is not its half-power point — at order 5 with half a decibel of ripple, the half-power point is six per cent higher. Bessel’s classical normalisation is for unit delay, which puts its half-power point somewhere else again.
Comparing the three as they come out of their definitions therefore compares three curves drawn at three different frequencies, and the result flatters or damns each family for reasons that have nothing to do with its behaviour. It is a large part of why Bessel filters have a reputation for being much shallower than they are.
So here every family is rescaled until its own measured half-power point lands at the same frequency, and the measurement is a bisection on the solved response rather than a factor from a table. The assertion in the figure is that all three land within three parts in a thousand of the target, which is the resolution of the bisection rather than a tolerance on the design.
The numbers, at order five
With that done, the comparison means something:
| at | passband deviation | delay variation | |
|---|---|---|---|
| Butterworth | −47.7 dB | 0.44 dB | 48.0% |
| Chebyshev | −64.0 dB | 0.50 dB | 49.0% |
| Bessel | −28.3 dB | 1.88 dB | 0.06% |
Three things in that table are worth stopping on.
The Chebyshev’s advantage is 16 dB. That is a factor of six in amplitude, at three times the corner, for the same number of components — a substantial and entirely real benefit, and it is now a number rather than the word “steeper”.
The Bessel’s disadvantage is larger than its advantage looks. It is 19 dB behind the Butterworth in the stopband, which is a factor of nine. Anyone choosing it for its transient behaviour is paying a great deal for it, and should know how much.
The passband column does not say what the adjectives suggest. The Butterworth’s deviation is 0.44 dB and the Chebyshev’s is 0.50 dB — essentially the same. “Maximally flat” is a statement about derivatives at zero frequency, not about how flat the passband is overall, and by the second measure the rippled filter is no worse. The Bessel’s 1.88 dB, meanwhile, is the largest of the three: its magnitude sags gently across the whole band, which is the price of a linear phase and is not described by the word “flat” at all.
What changes with the order
The slider moves the order from two to eight, and the trade moves with it in a way worth reading off rather than deriving.
The far stopband slope is 20 n decibels per decade for every family — a hundred at order five, a hundred and forty at order seven — because all three are all-pole filters of the same order and far enough out nothing else matters. What differs is how quickly each reaches that slope, and that is what the measurement at three times the corner captures.
The gap between families widens with order. At order three the Chebyshev leads the Bessel by 14 dB; at order five by 36; at order seven by 62. The choice of family therefore matters more, not less, as the filter gets more complicated — which is the opposite of the usual intuition that a high-order filter approaches an ideal brick wall regardless of how it is made.
And the Bessel’s delay variation falls with order, from 2.3% at order three to 0.06% at five and below 0.01% at seven, while the other two stay in the region of fifty per cent. That is the family doing exactly what it was designed for, and it is the only column in the table that improves with complexity.
Where the poles sit, and why the shapes follow
The three responses in the figure look different because the three sets of poles are in different places, and the correspondence is direct enough to be worth spelling out.
Butterworth’s poles are evenly spaced on a semicircle. All of them are the same distance from the origin, so all have the same natural frequency, and the ones nearest the imaginary axis — the pair at the ends of the arc — have a quality factor that rises with order. At order five that pair has a quality factor of about 1.6; at order nine, about 2.9. That pair is what produces both the sharpness of the transition and the peak in group delay.
Chebyshev’s poles are on an ellipse squashed toward the imaginary axis. Every pole is therefore closer to the axis than the corresponding Butterworth pole, the extreme pair much more so — at order five with half a decibel of ripple its quality factor is about 8.8, five times the Butterworth’s. A pole with a quality factor of nearly nine is a sharp resonance, and the passband ripple is that resonance showing through.
Bessel’s poles are spread along a curve that keeps them away from the axis. No pair has a high quality factor at any order, which is why nothing about the response is sharp: not the transition, not the delay, not the step.
So the three families are one statement made three ways. Everything sharp in a filter is a pole near the imaginary axis, and everything a sharp filter costs is what such a pole does in time. The comparison in this essay is a measurement of that single trade at three points along it.
What the realisation assumes
The networks measured here separate their sections with ideal followers, and the assumption deserves naming because it is a choice with consequences rather than a simplification.
With followers between them, each second-order section’s response is its own, and the cascade’s poles are exactly the union of the sections’ poles. Without them — in a passive ladder, which is the usual realisation at radio frequencies — each section loads the one before it, every pole moves, and the element values have to come from a different synthesis that accounts for the loading.
Both realise the same response, so nothing in the measurements changes. What changes is sensitivity: how much the response moves when a component is not quite its nominal value. A cascade of isolated sections has each pole determined by two or three components, so a one per cent error in one component moves one pole by about one per cent. A doubly-terminated passive ladder has a remarkable property — at frequencies where it transfers maximum power, the response is stationary with respect to every element value, so first-order errors cancel — which is why passive ladders are used where components are the limiting factor.
That is a genuine advantage of a realisation the figures here do not use, and it is worth recording as a limitation of the comparison rather than left for a reader to discover.
What is deliberately missing
There is a fourth family in every textbook version of this comparison, and it is not here.
The elliptic filter places zeros on the imaginary axis as well as poles, giving it a stopband that ripples rather than descending monotonically and a skirt steeper than a Chebyshev of the same order. It is the right answer whenever the requirement is a hard transition and the stopband only has to be below something rather than getting steadily further below it.
It is left out for a reason worth stating rather than hidden. Placing its zeros requires Jacobi elliptic functions, and realising them requires a ladder with resonant branches rather than the cascade of simple sections used here — a different and considerably longer piece of machinery. It sits further along the same axis as Chebyshev, steeper still and worse still in delay, so the three families here bracket the trade without it.
That is an honest gap rather than an oversight, and it is recorded as one in the library that builds these filters. A collection that quietly drops the inconvenient case and does not say so is doing something different from one that says what it left out.
The order that is not an integer
One practical note about the slider, since it moves through integers and a specification rarely lands on one.
The order of an all-pole filter is the number of poles, so it is a whole number by construction. A requirement that needs 4.3 poles has to be met with five, and the resulting filter exceeds the specification — usually by a comfortable margin, since each additional pole is worth twenty decibels per decade in the stopband.
That granularity is coarser than it looks. Between order four and order five, the attenuation at three times the corner improves by about ten decibels for a Butterworth and fifteen for a Chebyshev. So the choice is rarely between meeting the specification and missing it; it is between one design that misses and another that overshoots substantially, and the overshoot is paid for in components, cost, and — as this field keeps returning to — delay.
Which is one more argument for measuring the third column. If order five overshoots the magnitude requirement anyway, the freedom that buys can be spent on a gentler family: a Bessel of order six may meet the same stopband requirement as a Chebyshev of order four while having a hundredth of the delay variation. That trade is available only to somebody who has both numbers, and it is invisible in a table of adjectives.
The two routes, and what they caught
Every filter here exists twice: as a set of poles computed from a definition, and as a netlist of resistors, inductors, capacitors and followers. The figure asserts that the response of the second matches the response of the first to about a part in 1014 across four decades.
That agreement is not automatic, and the ways it can fail are instructive. Pairing the wrong two poles as conjugates produces sections with complex element values, which the network refuses. Getting the quality factor of a section wrong by a factor of two — easy, since the definition involves both a factor of two and a square root — produces a response that is still a low-pass of the right order with roughly the right corner, and is a different filter. Omitting the normalisation step produces three perfectly correct filters drawn at three different frequencies, which is the published state of the art.
Only the first of those three is caught by the network refusing. The other two are caught by comparing routes, and they are exactly the mistakes that would otherwise survive into a figure that looked entirely convincing.
The two highest orders make the last point on their own, and it is the one a second-order comparison hides: the families do not merely differ, they diverge, and the divergence is what decides whether the choice was worth making.
Choosing, in the terms the measurements support
The point of measuring is to be able to state the choice without adjectives, and it comes out roughly as follows.
If the requirement is a specified attenuation at a specified frequency and nothing else matters, the Chebyshev wins by a margin that grows with order, and the ripple it costs is invisible on any plot that is not magnified forty times. If the signal has edges in it — a pulse, a data stream, a step — and the shape matters, the Bessel is the only one of the three whose delay is flat, and everything else about it is worse. The Butterworth is the compromise, and its real distinction is not its flatness but that it has no bad property at all: no ripple, no sag, no extreme delay variation, nothing that will surprise anybody.
Which is, in the end, why it is the default. Not because maximal flatness is usually the requirement, but because a filter with no distinguishing weakness is the right choice when the requirement has not been stated precisely — and most of the time it has not.
What the comparison is worth once the requirement is stated
The paragraph above is an argument about defaults, and this field’s later essays are what happens when the requirement is stated. Each of them turns the comparison above into a different currency, and in two of the four the ranking changes.
What a steep skirt costs puts the delay column against the skirt column and finds the exchange rate: the order buys attenuation at twenty decibels per decade per pole and no arrangement of components changes that, so what a family buys is how quickly the slope is reached — and the steepest of the three distorts delay eight hundred times more than the gentlest. Flat magnitude, unflat delay is the measurement underneath it, and it is the one the classical comparison omits entirely: group delay varies by fifty per cent across the passband of the two families everybody uses.
What the filter in front costs states one requirement — a stated attenuation by the frequency that folds back into a stated band — and finds that the choice is not a decibel or two of skirt but a factor in the clock: Bessel demands 3.53 times Nyquist, Butterworth 2.08, Chebyshev 1.53. Priced that way the Bessel is not a compromise, it is more than twice as expensive as the Chebyshev in the one component the whole system is built around.
And the zeros that buy an order adds the family that is not on this page and that beats all three on the same requirement — an elliptic reaching forty decibels at twice the corner with two poles where the steepest all-pole family needs five — while charging in a currency none of the three has: a stopband that stops falling, measured at −38.7 dB and overtaken by an ordinary Chebyshev two corners out.
So the honest reading of this essay is that it establishes the axes rather than the answer. Three families measured on one corner, with everything each is good and bad at made into a number — and the choice between them decided by which of those numbers the requirement actually names.
What the choice of family decides later
Choosing a family is the first decision in the field and four later essays are consequences of it. What a steep skirt costs prices the attenuation against the delay flatness, which is the trade this comparison exists to state. Flat magnitude, unflat delay is the defect the Chebyshev choice produces, and Flat delay, bought with more delay is what repairing it costs. The zeros that buy an order is the fourth family, whose finite transmission zeros need a section the other three do not. A ladder is not a cascade is the realisation decision that follows the family decision, and Where the behaviour is written down is where all three families are the same pair of numbers read differently.
Part 1 on filter families
One argument about Filter families, and one of 2 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 16.
What this makes readable
Essays that name this one as a prerequisite.
- A ladder is not a cascade
- Flat delay, bought with more delay
- The band that closes with the order
- The bandwidth noise sees
- The ceiling is not at the output
- The corner error a filter hides in its sections
- The corner that moved
- The inductor that is an amplifier
- The number that was wrong
- The Q the amplifier decides
- The ratio that does not walk to one
- The resistor the noise comes from
- The ripple that is a temperature
- The same filter a thousand times larger
- The selectivity that is not free
- The staircase on the way out
- The termination an even order cannot have
- The two resistors a ladder was designed between
- The word length that is not a threshold
- What actually fills a null
- What the filter in front costs
- Which section goes first
The objects named here
The third axis, after the field and the idea: the things themselves, and every essay that touches each one.
BesselButterworthChebyshevFilter orderNormalisation
- Several sections, and the band they buy chebyshev, filter order