Networks, and how a solve is checked

The tolerance that can only take away

At a doubly terminated ladder's ripple peak the first derivative of the magnitude with respect to every reactance is zero to ten digits, which the rung below measured and which says nothing about how much the response moves. This says how: every second derivative is negative, so of six hundred ladders built from one per cent components not one is above nominal, the mean has shifted rather than the spread having grown, and doubling the tolerance quadruples the damage instead of doubling it.

Assumes: Every derivative, and the one that is zero · The tolerance that is not on any part

Every derivative, and the one that is zero computed the first derivative of a network’s response with respect to every element value in two solves, and found something specific about a doubly terminated ladder: at the passband ripple peaks the derivative with respect to every reactance vanishes. The classical argument for it is one sentence — at a ripple peak the lossless ladder is delivering all the power its source has, so there is nowhere up to go — and the measurement confirmed it to ten digits.

A derivative that is zero says the response does not move to first order. It says nothing about how much the response moves, and a specification is written about how much.

At the ripple peak the first derivative is 10⁻¹⁰ and the second is not. computed by solving, not by drawing. The relative first and second derivatives of |H| with respect to each reactance at the lower passband ripple peak of a 5th-order Chebyshev. The bars are the curvature; the first derivatives, printed beside them, are all below 10⁻⁸ and are the rung below's result. Every curvature is negative — the response is at a maximum in every one of these directions at once, because a doubly terminated lossless ladder at a ripple peak is delivering all the power its source has and there is nowhere up to go. Away from the peak, at 726.4 Hz, the first derivatives are 0.20, 0.04, 0.28 and the curvature is beside the point.
Fig. 1 The first and second derivatives of the magnitude with respect to each reactance, at the ripple peak. The first are all below 10⁻⁸; the second are not, and every one of them is negative.

Two derivatives, and which quantity they are of

There is a trap in getting from the first rung to this one and it is worth naming before the numbers.

The adjoint identity differentiates the complex response — that is what two solves of a transposed matrix give, and it is the right object for a pole or a phase. A passband specification is about the magnitude, and at a ripple peak the complex derivative with respect to a reactance is emphatically not zero while the magnitude derivative is zero to ten digits. Reading one as the other is a pleasant way to be wrong for a long time, so the magnitude’s own derivatives are computed explicitly:

dHdp=Re(HˉH)H,d2Hdp2=Re(HˉH)+H2(dH/dp)2H\frac{d|H|}{dp} = \frac{\mathrm{Re}(\bar H H')}{|H|}, \qquad \frac{d^2|H|}{dp^2} = \frac{\mathrm{Re}(\bar H H'') + |H'|^2 - (d|H|/dp)^2}{|H|}

Both are used in relative form, so that a fractional component change goes in and a fractional change in |H| comes out — which is the expansion the whole rung is about.

The cheap route, and the expensive one it is checked against

The second derivative is a central difference of the adjoint gradient: two adjoint runs per parameter, four solves each, against the four full re-solves per parameter a difference of differences on the response would need. The step is 10⁻⁴ rather than the first derivative’s 10⁻⁶, because a difference of a difference loses twice as many digits and the optimum step for a second derivative is the cube root of the arithmetic’s precision rather than its square root.

The expensive route is computed anyway and the two agree to better than one part in ten thousand on every element. That is what makes the cheap route evidence rather than a definition: it shares no arithmetic with the three-point second difference of the response — no transposed solve, no quadratic form, no derivative identity — so an agreement between them is a fact about the network.

Where one pole goes when each component is 1% high. computed by solving, not by drawing. A series R–L–C, its poles recovered by rooting the determinant, and the derivative of the upper one with respect to each element taken exactly from the two null vectors at the pole. The dashed lines are the first-order prediction for a 1 per cent change; the filled circles are where the root actually goes when the element is changed and the determinant re-rooted. The three directions are the argument: the resistance moves the pole along a circle of constant radius, because the natural frequency does not contain it — its normalised sensitivity has a real part of 3.1e-16. The inductance and the capacitance each carry exactly −½ of the radius, and imaginary parts that are exact negatives. At 1 per cent the prediction is out by 0.005 per cent of the pole's own magnitude.
Fig. 2 The same adjoint machinery pointed at a different object, from the rung below: the derivative of a pole rather than of a response, which is one negation away in the same quadratic form.

Every curvature is negative

At the lower ripple peak of a fifth-order 0.5 dB Chebyshev, the relative second derivatives with respect to the five reactances are −0.20, −0.25, −0.25, −0.53 and −0.53. All negative, with no exceptions and no small ones.

That is not a coincidence and it is the classical argument again, taken one derivative further. If the response is at the maximum it can possibly reach — every watt the source can deliver, arriving at the load — then it is at a maximum with respect to every direction in parameter space at once, and a maximum has a negative-definite second form. Any change to any reactance, in either direction, reduces it.

The two terminating resistors are the exception and they matter. Their first derivatives are ±0.5, not zero: the stationarity theorem is about the reactances, and a ladder is stationary because the match is perfect at that frequency, which the terminations define rather than achieve. A one per cent error in a termination is a first-order error and behaves like an error anywhere else.

The Butterworth ladder at order 5, expanded rather than looked up. computed by solving, not by drawing. The element values are the successive quotients of a continued-fraction expansion of (E+F)/(E−F), where E is the filter's own pole polynomial and F carries its reflection zeros. They are the numbers in every filter design table, and they agree with the closed form 2·sin((2k−1)π/2n) to twelve digits. The termination is 1.0000 times the source resistance, as an odd-order design must be.
Fig. 3 The element values the synthesis produces, terminations included. The two resistors are what make the ladder stationary and are the two components the stationarity does not protect.

That asymmetry is worth carrying: the property is bought by matching, and the parts that do the matching are exempt from the property they buy. A ladder built with one per cent reactances and five per cent terminations has its worst sensitivity in the two components a designer is least inclined to worry about — a reading the two resistors a ladder was designed between arrives at from the other direction.

Six hundred ladders

A statement about derivatives is a statement about a limit. What a yield is written about is a population, so six hundred ladders are built with every reactance drawn independently from ±1 per cent and measured at two frequencies.

At a stationary point a tolerance has a mean, not a spread. computed by solving, not by drawing. Six hundred ladders with every reactance drawn independently from ±1 per cent, measured at the ripple peak and at a frequency between the peaks. Away from the peak the distribution is centred on nominal and 325 of 600 are above it. At the peak none is: the whole distribution lies below, with a mean of -0.0030 per cent and a worst case of -0.0140. A yield calculation that assumes a symmetric spread at a frequency that has a stationary point is wrong in both directions at once — it allows parts above a limit that cannot exist, and it misses that the whole batch has moved.
Fig. 4 The distribution at the ripple peak and between the peaks, from the same six hundred draws. One of them is centred on nominal and the other is entirely below it.

Between the peaks, 325 of 600 are above nominal and the distribution is the symmetric spread every tolerance analysis assumes. At the peak, none is. Not one of six hundred is above nominal; the mean has moved to −0.0030 per cent and the worst case is −0.0140.

A yield calculation that assumes symmetry at a frequency that has a stationary point is wrong in two directions at once. It allows parts above a limit that cannot exist — so it over-estimates the failure rate at an upper limit and under-estimates it at a lower one — and it treats the whole batch’s shift as though it were noise.

Doubling the tolerance quadruples the damage

The order is read off directly by repeating the draw at four tolerances.

Doubling the tolerance doubles one spread and quadruples the other. computed by solving, not by drawing. The same four hundred draws at four tolerances. Between the peaks the spread is proportional to the tolerance — fitted exponent 1.000 — which is the ordinary first-order behaviour every tolerance analysis assumes. At the ripple peak the exponent is 1.999: the first-order term is not small there, it is absent, and what is left is the curvature. Tightening a tolerance from two per cent to one buys a factor of two between the peaks and a factor of four at them.
Fig. 5 The spread at each of four tolerances, at the peak and between the peaks. Fitted exponents 2.00 and 1.00.

Between the peaks the fitted exponent is 1.000: the spread is proportional to the tolerance, which is the ordinary behaviour. At the ripple peak it is 1.999. The first-order term is not small there, it is absent, and what is left is the curvature.

The practical reading is a cost argument. Tightening a tolerance from two per cent to one per cent buys a factor of two between the peaks and a factor of four at them — so the same expenditure is worth twice as much at exactly the frequencies where the specification is tightest, which is the opposite of the usual intuition that a stationary point is somewhere one can afford to be careless.

A distribution with a mean, elsewhere in this collection

A tolerance that shifts a batch rather than spreading it is not unique to this network, and the two other places it appears are worth putting beside it because the mechanism is different each time.

The error that is a distribution measured three hundred mismatched current mirrors and found the opposite arrangement: the mismatch does not move the mean at all and is the whole of the width, while a systematic error — the base currents — sets the mean and is not a distribution. There, the first-order term is present and the two effects are separable because one is deterministic.

At a stationary point a tolerance has a mean, not a spread. computed by solving, not by drawing. Six hundred ladders with every reactance drawn independently from ±5 per cent, measured at the ripple peak and at a frequency between the peaks. Away from the peak the distribution is centred on nominal and 316 of 600 are above it. At the peak none is: the whole distribution lies below, with a mean of -0.0742 per cent and a worst case of -0.3405. A yield calculation that assumes a symmetric spread at a frequency that has a stationary point is wrong in both directions at once — it allows parts above a limit that cannot exist, and it misses that the whole batch has moved.
Fig. 6 The same one-sided reading at five per cent rather than one or two. A distribution with a mean, elsewhere in this collection, is what a current mirror’s spread is — errors either side of a design value. This one has no mean to be spread about: every realisable component moves the response the same way, so the distribution has a wall at one end and a tail at the other.

The tolerance that is not on any part found a third arrangement: a quantity whose tolerance is not any component’s, because it is a ratio and the correlations decide it. That is a statement about the inputs rather than about the function they go through.

The three together make the point that a tolerance analysis has two halves — what varies, and what the network does with it — and that assuming a symmetric output from symmetric inputs is a claim about the second half rather than a property of the first.

Where the first rung’s result was doing its work

It is worth being precise about what the rung below did and did not establish, because the two results are often quoted as one.

That rung established that the ladder is better than a cascade — its first derivatives vanish where a cascade’s do not — and measured the poles’ own sensitivity beside it, finding a factor of 2.17 rather than the 10⁸ the magnitude’s stationarity might suggest. This rung establishes what the remaining sensitivity looks like: not a smaller spread of the same shape, but a different shape.

The two together say something a designer can act on. A ladder realisation converts a component tolerance into a small systematic loss at the ripple peaks and an ordinary spread elsewhere; a cascade converts it into an ordinary spread everywhere. If the specification is a ripple band, the ladder’s error eats into the band from one side only and the whole band can be allocated accordingly.

Which frequencies have this property

Not many, and that is the point. The stationarity holds at the ripple peaks of a doubly terminated lossless ladder — the frequencies where the reflection is exactly zero — and nowhere else. Between them the first derivatives are ordinary numbers and the distribution is ordinary.

At a stationary point a tolerance has a mean, not a spread. computed by solving, not by drawing. Six hundred ladders with every reactance drawn independently from ±2 per cent, measured at the ripple peak and at a frequency between the peaks. Away from the peak the distribution is centred on nominal and 325 of 600 are above it. At the peak none is: the whole distribution lies below, with a mean of -0.0119 per cent and a worst case of -0.0556. A yield calculation that assumes a symmetric spread at a frequency that has a stationary point is wrong in both directions at once — it allows parts above a limit that cannot exist, and it misses that the whole batch has moved.
Fig. 7 The same experiment at two per cent. The peak’s distribution is four times as wide as at one per cent and the middle frequency’s is twice as wide, which is the order result read off a histogram.

A higher-order design has more of these frequencies and they are more closely spaced, so the fraction of the passband that behaves this way rises with order — which is a further, unmeasured reason why the ladder realisation is preferred for high-order designs and why the preference is stated as folklore more often than as arithmetic.

What a trim can and cannot recover

A one-sided error with a known sign invites a correction, and it is worth saying exactly what is correctable.

The shift at a ripple peak is a loss of gain at that frequency, and the ladder’s stationarity means it is second order in the component errors. Scaling the whole response — a gain trim — recovers the loss at one frequency and moves the others, because the different peaks have different curvatures: in this design the upper peak’s curvatures are an order of magnitude larger than the lower peak’s, so a batch’s response sags more at the band edge than in the middle. The ripple has grown, and no scalar multiplies it back.

At the ripple peak the first derivative is 10⁻¹⁰ and the second is not. computed by solving, not by drawing. The relative first and second derivatives of |H| with respect to each reactance at the lower passband ripple peak of a 7th-order Chebyshev. The bars are the curvature; the first derivatives, printed beside them, are all below 10⁻⁸ and are the rung below's result. Every curvature is negative — the response is at a maximum in every one of these directions at once, because a doubly terminated lossless ladder at a ripple peak is delivering all the power its source has and there is nowhere up to go. Away from the peak, at 683.8 Hz, the first derivatives are 0.03, 0.25, 0.07 and the curvature is beside the point.
Fig. 8 The same curvature measurement on a seventh-order design. More peaks, larger curvatures at the band edge, and the same sign everywhere.

What would recover it is a trim on the terminations, because those are the components the stationarity does not cover and their first derivatives are ±0.5. A half per cent adjustment to a terminating resistor moves the passband by a quarter of a per cent at first order, which is far more than the second-order sag being corrected — so the adjustment exists and is delicate, which is a familiar combination.

Two things this does not settle

It is not a proof. Six hundred draws with none above nominal is strong evidence for a negative-definite form and it is not a demonstration that the form is negative-definite everywhere in the ±1 per cent box. The curvatures are measured at the nominal point; far enough away the expansion is not the function. What would settle it is the eigenvalues of the full Hessian rather than its diagonal, and the machinery here computes the off-diagonal terms as a by-product of the same two gradients — it is not exercised in this essay.

And loss breaks it. The argument requires the ladder to be lossless, because it turns on all the source’s power arriving at the load. Real inductors have a series resistance, some of that power is dissipated, and the response at the ripple peak is therefore not at its maximum — so the first derivative is not exactly zero and the one-sided result degrades continuously as the quality factor falls. The Q the components allow is where that resistance is measured; how far it has to rise before the distribution reopens is not measured anywhere.

The cross terms, and why the diagonal is not the whole story

The Hessian’s diagonal is what a per-component tolerance multiplies, and it is not what a batch experiences, because a batch varies every component at once.

Expanding to second order, the fractional change in |H| is the sum over pairs of 12Hijδiδj\tfrac12 H_{ij}\delta_i\delta_j. With independent component errors the off-diagonal terms have zero mean, so the mean shift is the diagonal alone — which is why the mean measured on six hundred draws matches the diagonal curvatures. The spread is not: it has the off-diagonal terms in it, and on a ladder they are not small, because neighbouring reactances of a ladder are strongly coupled by construction.

The machinery computes those pairs as a by-product: the gradient at p+h and at p−h contains the derivative with respect to every other parameter, so the mixed second derivatives cost nothing beyond what the diagonal already paid. What it would take to use them is a decision about what question they answer, and the honest position is that this essay measured the population directly rather than predicting it — six hundred solves being cheaper than the argument.

What a zero derivative was worth

The rung below found a derivative that vanishes and reported it as what it is: a first-order result. This one asks what the second order looks like and finds that the answer is not merely smaller. It has a different sign structure, a different scaling law, and a different consequence for a batch of parts — a mean that moves rather than a spread that grows.

Both came out of the same two solves of a transposed matrix, one differentiation apart, and both were checked against a route that shares no arithmetic with them.

The habit is worth stating in general because it recurs. A quantity reported as zero is a report about the leading term, and what is under it is a different object with a different scaling law and often a different sign structure. A sum that is exact makes the same point from the other side — an estimate that is exact in one term and a bound in the rest — and the cancellation that leaves a tail measures what survives a cancellation designed to be perfect. In all three the interesting arithmetic starts where the first-order description stops.

And in this one the answer to how much does it move turned out to have a different shape from the question. It does not move by a smaller amount in both directions. It moves in one direction, by a quantity that is quadratic in the tolerance, with a mean rather than a spread — what a network answers being, as usual, more specific than what was asked.

What six hundred samples could not have said

The six hundred ladders are what makes the direction claim convincing, and the rung above this one shows what they were not able to see.

The three tolerances that do nothing computes the whole second-derivative matrix rather than sampling it, and finds two of its five eigenvalues of order one and the other three nine decades down. That is a three-dimensional subspace of component variations along which the response does not move at all — and a sample of a five-dimensional box never lands on a three-dimensional subspace, so no number of built ladders would have found it. The six hundred here confirm a statement about a mean; the matrix there makes a statement about a structure.

The pairing is worth carrying because it says when to sample and when to differentiate. A sample answers what happens to a population, which is what a manufacturer needs, and it does so without any assumption that the perturbation is small. A derivative answers what the response is a function of, which is what a designer needs, and it is exact but local. The negative second derivatives on this page are what connect the two: they are why the population’s mean shifts rather than its spread growing, and they are computable in one solve by the method every derivative, and the one that is zero establishes.

And the first derivatives being zero is the reason any of this is interesting rather than routine. At a point where the gradient does not vanish, a tolerance moves the response first order in both directions and the mean does not shift at all. The one-sidedness measured here exists only at a stationary point — which is to say only at a doubly terminated ladder’s ripple peaks, and only while the terminations are the ones the design was made between.

That last clause is not a formality. The two resistors a ladder was designed between measures how narrow the condition is — a source outside 0.886 to 1.137 times the design resistance costs more than half a decibel — and everything on this page is computed inside it. A ladder driven from the wrong impedance is not stationary anywhere, so its response moves first order in every element and in both directions, and the one-sided quadratic behaviour measured here disappears entirely.

Part 3 on sensitivity

One argument about Sensitivity, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Adjoint networkComponent sensitivityComponent toleranceDevice mismatchLadder filterSecond-order approximationStationary pointYield