Networks, and how a solve is checked

Every derivative, and the one that is zero

How much does this response move if that capacitor is one per cent out? A difference quotient answers it one component at a time, in two solves each, and its best possible accuracy is four parts in a hundred million. Transposing the matrix and solving once more answers it for every component at once, exactly. Pointed at a claim this collection has made since its ladder essay and never tested directly — that a doubly terminated ladder's response is stationary in every element at its passband maxima — it returns two parts in ten billion, where the cascade realising the identical response returns 0.72.

Assumes: What a network answers, and how the answer is checked · The tolerance that is not on any part

Every design question in this collection that is not about a boundary is about a derivative. How much does the corner move if that capacitor is ten per cent low? Which of the eleven components in this filter is the one to buy the tight tolerance on? Is the response at this frequency stationary in the inductors, as the theory says it should be?

The obvious way to answer any of them is to move the component and solve again. It works, it needs no new ideas, and it has two costs that are usually treated as one. The first is arithmetic: a network with NN elements needs 2N2N extra solves for a decent difference quotient. The second is accuracy, and it is the one worth an essay, because a difference quotient of a solved circuit is a subtraction of two nearly equal numbers and inherits everything that implies.

There is a route with neither cost. It is one line of linear algebra, it has been known since the 1960s, and this collection has not had it.

One line

The network is Mx=bMx = b and the quantity wanted is one node voltage, which is eTxe^{\mathsf{T}}x for the unit vector ee at that node. Differentiate with respect to any parameter pp that appears only in the matrix:

Mxp=Mpxxoutp=eTM1MpxM\frac{\partial x}{\partial p} = -\frac{\partial M}{\partial p}x \qquad\Longrightarrow\qquad \frac{\partial x_\mathrm{out}}{\partial p} = -e^{\mathsf{T}}M^{-1}\frac{\partial M}{\partial p}x

The whole of the method is in what that expression does not depend on. The row eTM1e^{\mathsf{T}}M^{-1} is the same for every parameter. Solve MTy=eM^{\mathsf{T}}y = e once and every derivative in the network is yT(M/p)x-y^{\mathsf{T}}(\partial M/\partial p)x — an inner product over the two or four entries that element stamps, which costs nothing at all.

So: one forward solve for the circuit, one transposed solve, and then every derivative for free. Two solves, whatever the network holds.

M/p\partial M/\partial p is the element’s own stamp differentiated, which is why the code is short: a resistor’s stamp is an admittance so its derivative with respect to resistance carries the 1/R2-1/R^2 the chain rule gives; a capacitor’s is ss times the same pattern; an inductor’s own row carries sL-sL on the diagonal; a transconductance’s four entries are linear in gmg_m already.

One detail has to be right and is easy to get wrong in a way that hides. It is the transpose, not the conjugate transpose. The system is complex and the quantity is eTxe^{\mathsf{T}}x rather than an inner product, so conjugating gives an answer that is correct for a real network at direct current and wrong at every other frequency — which is a pleasant way to be wrong for a long time, since every check anybody writes first is at direct current.

Checking it against the route it replaces

The exact route and the difference quotient share nothing: one differentiates the matrix and the other re-solves a perturbed netlist. That makes them a genuine pair, and the comparison is the figure this essay opens with.

A difference quotient is best at a step of 1e-5, and is 1e+6 times worse at 10⁻¹¹. computed by solving, not by drawing. The worst disagreement between the adjoint network's derivatives and a central difference quotient of the same quantities, against the fractional step the quotient is taken with, on a 7-element ladder at 1000 Hz. The curve has a minimum because two errors pull opposite ways: the curvature the quotient neglects falls as the square of the step, and the digits its subtraction destroys rise as one over the step. The best it reaches is 1.4e-9, against the 2.3e-11 that the two-thirds power of the machine epsilon predicts. The exact route costs 2 solves against 15, and has neither error term.
Fig. 1 The worst disagreement between the two routes on a seven-element ladder, against the fractional step the difference quotient is taken with. The curve has a minimum, and the minimum is the best a difference quotient can do.

The curve is a V, and each side of it is a different failure.

On the right, at large steps, the quotient is measuring a chord rather than a tangent: the error is the curvature it neglects and falls as the square of the step. On the left, at small steps, the two solved voltages agree in their leading digits and the subtraction throws those digits away: the error is the rounding in what is left, divided by a step that is getting smaller, so it rises as one over the step.

The sum has a minimum where the two are equal, at a step near ε1/3\varepsilon^{1/3}, and the best it can reach is about ε2/3\varepsilon^{2/3} — four parts in 101110^{11} for double precision. Measured on this network at a kilohertz: the best step is 10510^{-5}, it reaches 1.4×1091.4\times10^{-9}, and at 101110^{-11} it has degenerated to 1.4×1031.4\times10^{-3}. A million times worse for asking a more careful question.

That is the second cost of the difference quotient and it is the one nobody budgets for. The first — fifteen solves against two on this network — is arithmetic that gets cheaper every year. The second does not.

At a passband maximum the ladder's every element measures 1e-8 and the cascade's 0.74. computed by solving, not by drawing. The magnitude sensitivity of the response to each reactive element of a 7th-order Chebyshev, for a doubly terminated ladder and for a buffered cascade of the identical response, at 421.2 Hz — a passband maximum. Each is the exact derivative from the adjoint network, two solves for the whole list. The ladder's largest is 9.7e-9 and the cascade's 7.4e-1. What is stationary is the magnitude: the phase sensitivities are ordinary at the same frequency, and the same ladder 379 Hz away is ordinary too.
Fig. 2 The seventh-order pair, at a passband maximum of 421.2 Hz. The ladder’s largest magnitude sensitivity is 9.7×10⁻⁹ and the cascade’s 7.4×10⁻¹ — eight orders of magnitude, from two solves rather than from fourteen perturbed ones. What is stationary is the magnitude alone: the phase sensitivities at the same frequency are ordinary, and so is the same ladder 379 Hz away.

Where it points, which is at a claim this site has been making

The filters field has said since its ladder essay that a doubly terminated LC ladder is stationary in its components at the frequencies where its passband response reaches a maximum. The argument is an energy argument and it is short: at those frequencies a lossless network between fixed terminations is delivering exactly the power the source has available, so any change in any component in any direction can only take some away. The first derivative of the response with respect to every reactive element is therefore zero there.

That has been tested by perturbing every component by one per cent, perturbing it by ten, and asking whether the deviation grew by ten or by a hundred. The measured exponents are 0.99 for the cascade and 2.00 for the ladder, and they are a good test of an exponent.

They are an indirect test of the statement, which is about a derivative. Here is the derivative.

At a passband maximum the ladder's every element measures 2e-10 and the cascade's 0.72. computed by solving, not by drawing. The magnitude sensitivity of the response to each reactive element of a 5th-order Chebyshev, for a doubly terminated ladder and for a buffered cascade of the identical response, at 554.9 Hz — a passband maximum. Each is the exact derivative from the adjoint network, two solves for the whole list. The ladder's largest is 2.3e-10 and the cascade's 7.2e-1. What is stationary is the magnitude: the phase sensitivities are ordinary at the same frequency, and the same ladder 499 Hz away is ordinary too.
Fig. 3 Every reactive element’s magnitude sensitivity, for a fifth-order Chebyshev realised as a doubly terminated ladder and as a buffered cascade of the identical response, at the lower of the two ripple peaks. One pair of marks per element, on an axis spanning thirteen decades.

At the lower ripple peak of a fifth-order Chebyshev, the ladder’s five reactive elements have magnitude sensitivities of 1.8×10101.8\times10^{-10}, 2.3×10102.3\times10^{-10}, 2.1×1010-2.1\times10^{-10}, 2.3×10102.3\times10^{-10} and 1.8×10101.8\times10^{-10}. The cascade realising the identical response, measured at the identical frequency, has 0.48, 0.45, 0.33, −0.54 and −0.72.

Nine orders of magnitude, element by element, in two solves each.

Three things make that number mean what it looks like it means, and each is a separate measurement.

The same elements are ordinary elsewhere. Ninety per cent of the way to the peak — 500 Hz against 555 — the same ladder’s sensitivities are 0.029, 0.056, −0.048, 0.056, 0.029. So what is being measured is the stationarity of the response, not some peculiarity of an LC ladder’s components.

Off the maximum the two realisations are within 12× of each other. computed by solving, not by drawing. The magnitude sensitivity of the response to each reactive element of a 5th-order Chebyshev, for a doubly terminated ladder and for a buffered cascade of the identical response, at 499.4 Hz — a frequency that is not a maximum. Each is the exact derivative from the adjoint network, two solves for the whole list. The ladder's largest is 5.6e-2 and the cascade's 6.8e-1. What is stationary is the magnitude: the phase sensitivities are ordinary at the same frequency, and the same ladder 499 Hz away is ordinary too.
Fig. 4 The same two realisations a tenth of the way off the maximum. The ladder is better by a factor of about twelve, which is a realisation being better; at the maximum it is better by a factor of 10910^9, which is a derivative being zero.

Only the magnitude is stationary. A relative sensitivity is a complex number: its real part is the fractional change in magnitude and its imaginary part the change in phase. The real parts at the peak are 101010^{-10}; the imaginary parts are of order one. The ladder’s phase moves with its components exactly as much as anything else’s does, which the energy argument never claimed otherwise and which nobody states.

The residual is a measurement of the peak, not of the ladder. Near a maximum the derivative vanishes linearly in the distance from it, so the 2×10102\times10^{-10} left over is reporting how precisely the golden-section search located the peak. At the upper ripple peak the same ladder returns 8×1088\times10^{-8} — four hundred times larger, from a peak located four hundred times less well.

One per cent on one component, in two realisations of the same order-5 filtercomputed by solving, not by drawing. The two realisations agree to 3e-14 dB before anything is moved. Moving each element in turn by 1%, the worst deviation at the 2 ripple peaks is 0.1621 dB for the cascade and 0.0030 dB for the LC ladder — 54 times smaller. Across the whole passband the two are within 3% of each other, because near the band edge both are dominated by the response's own steepness rather than by the realisation.-1.50-1-0.50000.50002004006008001e+3frequency (hertz)passband (decibels)nominal, both realisationscascade, F0C up 1%ladder, L2 up 1%ripple peaks — where the ladder is stationarysolved, then checked — one design, two realisations54× less movement at the peaks
Fig. 5 The rung this confirms, from the filters field: the same two realisations with every component perturbed by a stated fraction, and the exponents fitted over two decades of tolerance — 0.99 against 2.00. Drag it through the tolerance.
At a passband maximum the ladder's every element measures 2e-9 and the cascade's 0.75. computed by solving, not by drawing. The magnitude sensitivity of the response to each reactive element of a 9th-order Chebyshev, for a doubly terminated ladder and for a buffered cascade of the identical response, at 335.9 Hz — a passband maximum. Each is the exact derivative from the adjoint network, two solves for the whole list. The ladder's largest is 1.6e-9 and the cascade's 7.5e-1. What is stationary is the magnitude: the phase sensitivities are ordinary at the same frequency, and the same ladder 302 Hz away is ordinary too.
Fig. 6 The ninth-order pair. The ladder’s nine reactive elements are still at 10910^{-9} at the peak and the cascade’s are still of order one, so the stationarity is a property of the structure rather than something that gives out as the network grows.

The two elements that are not stationary, and the number they return

The list the adjoint returns has seven entries on a fifth-order ladder, and the section above quoted five of them. The other two are the terminating resistors, and they are the sharpest result here.

At the same ripple peak, at the same frequency, in the same solve:

element magnitude sensitivity phase sensitivity
Rs −0.500000 −1.8×10⁻¹⁰
L0 1.8×10⁻¹⁰ −0.501
C1 2.3×10⁻¹⁰ −0.725
L2 −2.1×10⁻¹⁰ −0.447
C3 2.3×10⁻¹⁰ −0.725
L4 1.8×10⁻¹⁰ −0.501
Rl +0.500000 −1.8×10⁻¹⁰

The table is two statements and each is exact.

The energy argument is about a lossless two-port between fixed terminations, and it says nothing whatever about the terminations. So the reactive elements are stationary and the resistors are not, and what the resistors return is not merely non-zero but exactly a half — because at a passband maximum the network is delivering the available power, the transfer is the matched half, and a fractional change in either resistance moves that by half of itself, with opposite signs. The theorem’s exclusion is as measurable as the theorem.

And the two columns are complements. The reactive elements have no magnitude sensitivity and ordinary phase sensitivity; the resistors have exactly the reverse. Nothing in the derivation predicts the second half of that, and it falls out of the same two solves.

The symmetry is a free check on top: L0 and L4 return the same number to twelve figures, as do C1 and C3, because the ladder is symmetric and nothing in the arithmetic knows it.

What a sensitivity is worth to somebody buying components

The reason to have every derivative rather than one is that the question a designer asks is about the whole network at once.

A relative sensitivity S=(p/v)(v/p)S = (p/v)(\partial v/\partial p) is dimensionless and is what a tolerance multiplies: a component tt per cent out moves the response by StS \cdot t per cent, to first order. So the worst case over independent tolerances is Siti\sum|S_i|t_i, and the list of Si|S_i| is a shopping list ordered by what it is worth paying for.

It also answers a question a Monte Carlo cannot answer cheaply: which component, if any, is the one to tighten. Two thousand draws give a spread; they do not attribute it. The derivatives attribute it exactly and cost two solves.

A divider of 4 equal 1.0% resistors, solved 3000 times. computed by solving, not by drawing. Every resistor drawn from its tolerance band and the divider solved, 3000 times. The worst case is ±1.000% — the part tolerance itself, and it does not improve when the divider is built from more parts — while the measured spread is 0.2944% and the worst of 3000 draws reached 82% of the bound. The root-sum-square, offered as though it were a standard deviation, is 1.70 of one here: for uniformly distributed parts it is √3 σ, a coverage of about 92%.
Fig. 7 The statistical route from this field’s own tolerance essay, and the one the derivatives complement rather than replace: a spread that is narrower than the worst case by a factor that depends on how many components there are.

Neither route replaces the other, and the reason is in the word first order. The derivative is exact and it is a derivative; a ten per cent capacitor is not an infinitesimal one, and on the ladder above the whole point is that the first order vanishes and what is left is second. A sensitivity list of zeros is a correct answer that tells a manufacturer nothing about how bad a one per cent ladder is — which is the measurement the rung this confirms exists to make.

Two more things the transposed solve knows

The adjoint solve has an interpretation, and it is one this collection has already met.

MTy=eM^{\mathsf{T}}y = e is the network driven backwards: a unit current injected at the output node, with every source turned off, and yy the voltages that result. For a network of two-terminal elements MM is symmetric, so the transposed network is the network, and the adjoint solve is the same circuit excited at its output — which is exactly the statement of reciprocity this field measures in its own essay on it.

Put a transconductance in and MM stops being symmetric, reciprocity fails, and the adjoint network stops being the network — it becomes the network with every controlled source reversed, source for sensing node. The method does not care: it transposes the matrix it has. But it is worth seeing that the two facts are one fact, because it says where the intuition “drive it backwards” stops being available and the arithmetic has to be trusted instead.

The second thing is a limitation and it is sharp. This computes the derivative of one output with respect to every parameter. The transposed solve is per output, not per parameter. A design question about two outputs needs three solves, and about a hundred outputs needs a hundred and one — at which point the difference quotient, which is per parameter and independent of how many outputs are read, is the cheaper route again. Which of the two is right is a question about the shape of the problem, and the shape here — one response, many components — is the one the adjoint was invented for.

What the derivatives are used for

One transposed solve gives every derivative, and this collection spends them in four places. A ladder is not a cascade is the claim they confirm, turning a comparison of two realisations into an exact statement about stationarity. The tolerance that is not on any part is the statistical route they replace — three thousand draws give a spread and attribute nothing. The derivative of a root carries the same machinery to the poles rather than to the response. The matrix that is ill, and the answer that is not is the reminder that an exact derivative of a badly conditioned answer is still a derivative of that answer, and Two solves that add, and the one that does not is the other place where two solves say more than one.

What is checked

The two routes are required to agree, and the assertion is on the shape of the disagreement rather than on its size: the difference quotient must be worse at a large step, worse at a small one, and best somewhere in between at a value near the two-thirds power of the machine epsilon. A check that merely required agreement at one step would pass with a wrong reference and a lucky step.

The exact route is checked against a closed form where one exists — a loaded divider, whose derivative with respect to each resistance is two lines of algebra — so that the pair above is anchored to something neither of them computed.

The stationarity result carries four assertions and only the first is the headline. Every reactive element of the ladder is asserted to be under 10610^{-6} at the peak; the same elements are asserted to be four orders larger a tenth of the way off it; the cascade is asserted to be six orders larger at the same frequency; and the ladder’s imaginary parts are asserted to be six orders larger than its real ones, which is the claim that what is stationary is the magnitude alone.

The direct-current case is its own check: at zero frequency an inductor is a short and a capacitor an open, so the derivative of any node voltage with respect to either is identically zero, and the method returns exactly zero rather than a small number.

What is not computed: second derivatives, which is what the ladder’s own argument actually needs and which would need one solve per parameter again; the sensitivity of a pole rather than of a response, which is a derivative of a root and is a different object; and anything about a nonlinear network, where the matrix is the Jacobian at an operating point and the same line applies to the linearised system and to nothing else.

What the transpose is, and what it costs to have

The one solve that answers for every element is not a trick of the arithmetic; it is the adjoint network, which is the same network with its excitation and its reading exchanged. For a reciprocal network that is the network itself, which is why the extra solve is a solve of the same matrix transposed rather than of a new one — the reading that does not care which way round it is is where reciprocity is measured, at five parts in 101410^{14} over four decades, and it is the property this whole method stands on. Put a transconductance of eight tenths of a femtosiemens anywhere in the network and the two readings part company; the transposed solve still returns derivatives, and they are the derivatives of the adjoint rather than of the circuit.

Three results the method made available

The exactness matters less than the completeness, and three later measurements are things a difference quotient could not have produced.

The tolerance that can only take away needed every second derivative at a ripple peak, found them all negative, and could then say that of six hundred ladders built from one per cent components not one is above nominal — a statement about a direction rather than a magnitude, which follows from the signs and from nothing else.

The three tolerances that do nothing needed the whole matrix rather than its diagonal, and found two of five eigenvalues of order one with the other three nine decades down — a three-dimensional subspace of component variations that do nothing at all, which a sample of a five-dimensional box never lands on and which six hundred built ladders could not have found.

And the derivative of a root is the case where the method has to be replaced rather than reused: a pole is a value of ss at which the matrix loses rank, not a response at a fixed frequency, so its derivative comes from two null vectors and a division. Pointed at the two realisations this page compares, it puts the ladder ahead by a factor of 2.17 rather than by eight orders — because what is stationary is the magnitude at one frequency, and that says nothing about where the poles are.

The three together are the argument for having built the method rather than a faster difference quotient. Speed was not the reason: on a network of this size the two routes take comparable time. The reason is that a complete gradient is an object with structure — signs, eigenvalues, a null space — and none of those is visible one component at a time, however accurately each component is measured.

The exactness is worth something on its own as well, and it is worth being precise about what. Four parts in a hundred million is an excellent difference quotient and it is a bound rather than an error: the step size that achieves it was found by search, and a smaller or larger one does worse. An adjoint solve has no step size, so there is nothing to tune and nothing that degrades on a badly-scaled network — which matters most on exactly the networks a difference quotient is worst on.

Part 1 on sensitivity

One argument about Sensitivity, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 21.

What this makes readable

Essays that name this one as a prerequisite.

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Adjoint networkComponent sensitivityComponent toleranceModel rangeNumerical errorRealisationReciprocityVerification