Every derivative, and the one that is zero
Assumes: What a network answers, and how the answer is checked · The tolerance that is not on any part
Every design question in this collection that is not about a boundary is about a derivative. How much does the corner move if that capacitor is ten per cent low? Which of the eleven components in this filter is the one to buy the tight tolerance on? Is the response at this frequency stationary in the inductors, as the theory says it should be?
The obvious way to answer any of them is to move the component and solve again. It works, it needs no new ideas, and it has two costs that are usually treated as one. The first is arithmetic: a network with elements needs extra solves for a decent difference quotient. The second is accuracy, and it is the one worth an essay, because a difference quotient of a solved circuit is a subtraction of two nearly equal numbers and inherits everything that implies.
There is a route with neither cost. It is one line of linear algebra, it has been known since the 1960s, and this collection has not had it.
One line
The network is and the quantity wanted is one node voltage, which is for the unit vector at that node. Differentiate with respect to any parameter that appears only in the matrix:
The whole of the method is in what that expression does not depend on. The row is the same for every parameter. Solve once and every derivative in the network is — an inner product over the two or four entries that element stamps, which costs nothing at all.
So: one forward solve for the circuit, one transposed solve, and then every derivative for free. Two solves, whatever the network holds.
is the element’s own stamp differentiated, which is why the code is short: a resistor’s stamp is an admittance so its derivative with respect to resistance carries the the chain rule gives; a capacitor’s is times the same pattern; an inductor’s own row carries on the diagonal; a transconductance’s four entries are linear in already.
One detail has to be right and is easy to get wrong in a way that hides. It is the transpose, not the conjugate transpose. The system is complex and the quantity is rather than an inner product, so conjugating gives an answer that is correct for a real network at direct current and wrong at every other frequency — which is a pleasant way to be wrong for a long time, since every check anybody writes first is at direct current.
Checking it against the route it replaces
The exact route and the difference quotient share nothing: one differentiates the matrix and the other re-solves a perturbed netlist. That makes them a genuine pair, and the comparison is the figure this essay opens with.
The curve is a V, and each side of it is a different failure.
On the right, at large steps, the quotient is measuring a chord rather than a tangent: the error is the curvature it neglects and falls as the square of the step. On the left, at small steps, the two solved voltages agree in their leading digits and the subtraction throws those digits away: the error is the rounding in what is left, divided by a step that is getting smaller, so it rises as one over the step.
The sum has a minimum where the two are equal, at a step near , and the best it can reach is about — four parts in for double precision. Measured on this network at a kilohertz: the best step is , it reaches , and at it has degenerated to . A million times worse for asking a more careful question.
That is the second cost of the difference quotient and it is the one nobody budgets for. The first — fifteen solves against two on this network — is arithmetic that gets cheaper every year. The second does not.
Where it points, which is at a claim this site has been making
The filters field has said since its ladder essay that a doubly terminated LC ladder is stationary in its components at the frequencies where its passband response reaches a maximum. The argument is an energy argument and it is short: at those frequencies a lossless network between fixed terminations is delivering exactly the power the source has available, so any change in any component in any direction can only take some away. The first derivative of the response with respect to every reactive element is therefore zero there.
That has been tested by perturbing every component by one per cent, perturbing it by ten, and asking whether the deviation grew by ten or by a hundred. The measured exponents are 0.99 for the cascade and 2.00 for the ladder, and they are a good test of an exponent.
They are an indirect test of the statement, which is about a derivative. Here is the derivative.
At the lower ripple peak of a fifth-order Chebyshev, the ladder’s five reactive elements have magnitude sensitivities of , , , and . The cascade realising the identical response, measured at the identical frequency, has 0.48, 0.45, 0.33, −0.54 and −0.72.
Nine orders of magnitude, element by element, in two solves each.
Three things make that number mean what it looks like it means, and each is a separate measurement.
The same elements are ordinary elsewhere. Ninety per cent of the way to the peak — 500 Hz against 555 — the same ladder’s sensitivities are 0.029, 0.056, −0.048, 0.056, 0.029. So what is being measured is the stationarity of the response, not some peculiarity of an LC ladder’s components.
Only the magnitude is stationary. A relative sensitivity is a complex number: its real part is the fractional change in magnitude and its imaginary part the change in phase. The real parts at the peak are ; the imaginary parts are of order one. The ladder’s phase moves with its components exactly as much as anything else’s does, which the energy argument never claimed otherwise and which nobody states.
The residual is a measurement of the peak, not of the ladder. Near a maximum the derivative vanishes linearly in the distance from it, so the left over is reporting how precisely the golden-section search located the peak. At the upper ripple peak the same ladder returns — four hundred times larger, from a peak located four hundred times less well.
The two elements that are not stationary, and the number they return
The list the adjoint returns has seven entries on a fifth-order ladder, and the section above quoted five of them. The other two are the terminating resistors, and they are the sharpest result here.
At the same ripple peak, at the same frequency, in the same solve:
| element | magnitude sensitivity | phase sensitivity |
|---|---|---|
| Rs | −0.500000 | −1.8×10⁻¹⁰ |
| L0 | 1.8×10⁻¹⁰ | −0.501 |
| C1 | 2.3×10⁻¹⁰ | −0.725 |
| L2 | −2.1×10⁻¹⁰ | −0.447 |
| C3 | 2.3×10⁻¹⁰ | −0.725 |
| L4 | 1.8×10⁻¹⁰ | −0.501 |
| Rl | +0.500000 | −1.8×10⁻¹⁰ |
The table is two statements and each is exact.
The energy argument is about a lossless two-port between fixed terminations, and it says nothing whatever about the terminations. So the reactive elements are stationary and the resistors are not, and what the resistors return is not merely non-zero but exactly a half — because at a passband maximum the network is delivering the available power, the transfer is the matched half, and a fractional change in either resistance moves that by half of itself, with opposite signs. The theorem’s exclusion is as measurable as the theorem.
And the two columns are complements. The reactive elements have no magnitude sensitivity and ordinary phase sensitivity; the resistors have exactly the reverse. Nothing in the derivation predicts the second half of that, and it falls out of the same two solves.
The symmetry is a free check on top: L0 and L4 return the same number to twelve figures, as do C1 and C3, because the ladder is symmetric and nothing in the arithmetic knows it.
What a sensitivity is worth to somebody buying components
The reason to have every derivative rather than one is that the question a designer asks is about the whole network at once.
A relative sensitivity is dimensionless and is what a tolerance multiplies: a component per cent out moves the response by per cent, to first order. So the worst case over independent tolerances is , and the list of is a shopping list ordered by what it is worth paying for.
It also answers a question a Monte Carlo cannot answer cheaply: which component, if any, is the one to tighten. Two thousand draws give a spread; they do not attribute it. The derivatives attribute it exactly and cost two solves.
Neither route replaces the other, and the reason is in the word first order. The derivative is exact and it is a derivative; a ten per cent capacitor is not an infinitesimal one, and on the ladder above the whole point is that the first order vanishes and what is left is second. A sensitivity list of zeros is a correct answer that tells a manufacturer nothing about how bad a one per cent ladder is — which is the measurement the rung this confirms exists to make.
Two more things the transposed solve knows
The adjoint solve has an interpretation, and it is one this collection has already met.
is the network driven backwards: a unit current injected at the output node, with every source turned off, and the voltages that result. For a network of two-terminal elements is symmetric, so the transposed network is the network, and the adjoint solve is the same circuit excited at its output — which is exactly the statement of reciprocity this field measures in its own essay on it.
Put a transconductance in and stops being symmetric, reciprocity fails, and the adjoint network stops being the network — it becomes the network with every controlled source reversed, source for sensing node. The method does not care: it transposes the matrix it has. But it is worth seeing that the two facts are one fact, because it says where the intuition “drive it backwards” stops being available and the arithmetic has to be trusted instead.
The second thing is a limitation and it is sharp. This computes the derivative of one output with respect to every parameter. The transposed solve is per output, not per parameter. A design question about two outputs needs three solves, and about a hundred outputs needs a hundred and one — at which point the difference quotient, which is per parameter and independent of how many outputs are read, is the cheaper route again. Which of the two is right is a question about the shape of the problem, and the shape here — one response, many components — is the one the adjoint was invented for.
What the derivatives are used for
One transposed solve gives every derivative, and this collection spends them in four places. A ladder is not a cascade is the claim they confirm, turning a comparison of two realisations into an exact statement about stationarity. The tolerance that is not on any part is the statistical route they replace — three thousand draws give a spread and attribute nothing. The derivative of a root carries the same machinery to the poles rather than to the response. The matrix that is ill, and the answer that is not is the reminder that an exact derivative of a badly conditioned answer is still a derivative of that answer, and Two solves that add, and the one that does not is the other place where two solves say more than one.
What is checked
The two routes are required to agree, and the assertion is on the shape of the disagreement rather than on its size: the difference quotient must be worse at a large step, worse at a small one, and best somewhere in between at a value near the two-thirds power of the machine epsilon. A check that merely required agreement at one step would pass with a wrong reference and a lucky step.
The exact route is checked against a closed form where one exists — a loaded divider, whose derivative with respect to each resistance is two lines of algebra — so that the pair above is anchored to something neither of them computed.
The stationarity result carries four assertions and only the first is the headline. Every reactive element of the ladder is asserted to be under at the peak; the same elements are asserted to be four orders larger a tenth of the way off it; the cascade is asserted to be six orders larger at the same frequency; and the ladder’s imaginary parts are asserted to be six orders larger than its real ones, which is the claim that what is stationary is the magnitude alone.
The direct-current case is its own check: at zero frequency an inductor is a short and a capacitor an open, so the derivative of any node voltage with respect to either is identically zero, and the method returns exactly zero rather than a small number.
What is not computed: second derivatives, which is what the ladder’s own argument actually needs and which would need one solve per parameter again; the sensitivity of a pole rather than of a response, which is a derivative of a root and is a different object; and anything about a nonlinear network, where the matrix is the Jacobian at an operating point and the same line applies to the linearised system and to nothing else.
What the transpose is, and what it costs to have
The one solve that answers for every element is not a trick of the arithmetic; it is the adjoint network, which is the same network with its excitation and its reading exchanged. For a reciprocal network that is the network itself, which is why the extra solve is a solve of the same matrix transposed rather than of a new one — the reading that does not care which way round it is is where reciprocity is measured, at five parts in over four decades, and it is the property this whole method stands on. Put a transconductance of eight tenths of a femtosiemens anywhere in the network and the two readings part company; the transposed solve still returns derivatives, and they are the derivatives of the adjoint rather than of the circuit.
Three results the method made available
The exactness matters less than the completeness, and three later measurements are things a difference quotient could not have produced.
The tolerance that can only take away needed every second derivative at a ripple peak, found them all negative, and could then say that of six hundred ladders built from one per cent components not one is above nominal — a statement about a direction rather than a magnitude, which follows from the signs and from nothing else.
The three tolerances that do nothing needed the whole matrix rather than its diagonal, and found two of five eigenvalues of order one with the other three nine decades down — a three-dimensional subspace of component variations that do nothing at all, which a sample of a five-dimensional box never lands on and which six hundred built ladders could not have found.
And the derivative of a root is the case where the method has to be replaced rather than reused: a pole is a value of at which the matrix loses rank, not a response at a fixed frequency, so its derivative comes from two null vectors and a division. Pointed at the two realisations this page compares, it puts the ladder ahead by a factor of 2.17 rather than by eight orders — because what is stationary is the magnitude at one frequency, and that says nothing about where the poles are.
The three together are the argument for having built the method rather than a faster difference quotient. Speed was not the reason: on a network of this size the two routes take comparable time. The reason is that a complete gradient is an object with structure — signs, eigenvalues, a null space — and none of those is visible one component at a time, however accurately each component is measured.
The exactness is worth something on its own as well, and it is worth being precise about what. Four parts in a hundred million is an excellent difference quotient and it is a bound rather than an error: the step size that achieves it was found by search, and a smaller or larger one does worse. An adjoint solve has no step size, so there is nothing to tune and nothing that degrades on a badly-scaled network — which matters most on exactly the networks a difference quotient is worst on.
Part 1 on sensitivity
One argument about Sensitivity, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 21.
What this makes readable
Essays that name this one as a prerequisite.
The objects named here
The third axis, after the field and the idea: the things themselves, and every essay that touches each one.
Adjoint networkComponent sensitivityComponent toleranceModel rangeNumerical errorRealisationReciprocityVerification
- Ten seconds, and fifteen minutes component tolerance, model range, verification
- The floor that outlives the arithmetic component sensitivity, numerical error, realisation
- The loop that never crosses model range, numerical error, verification
- The optimum a spectrum moves model range, numerical error, verification
- The reading a data sheet does not take component tolerance, model range, verification
- Exact outside and wrong within model range, verification