Before the steady state

Two ladders the terminals cannot tell apart

A thermal path drawn as a ladder and the same path drawn as a sum of exponentials are called different models of one object, and the difference between them has never been priced because pricing it needs an exact answer. Solved in closed form, the sum is 0.950 per cent high at worst and never low; the marched netlist is right to a part in 21,169; and the largest disagreement in the picture was 2.919 per cent that has nothing to do with heat at all, which reading the curve one sample differently removes.

Assumes: The diode that conducts backwards · Exact outside and wrong within · Where the behaviour is written down

The pulse the heatsink does not feel drew a junction’s single-pulse thermal impedance two ways. One was a netlist of three resistances and three heat capacities, marched. The other was the sum every account of a thermal path writes down — each stage’s own resistance times 1et/RC1 - e^{-t/RC}, with a time constant built from that stage’s own two components — and the essay called it an approximation, on the grounds that a ladder’s stages load each other: the die charges into a case that is itself moving.

One pulse from cold: the impedance is three plateaux, not a resistance. computed by solving, not by drawing. A hundred watts applied once, from ambient, and the junction's rise divided by it. The curve has three shoulders because the path has three stages: the die fills in about 2.4 ms, the case in 0.20 s, the heatsink in 8.0 minutes. Reading a junction temperature off the 9.7 K/W total is right only for pulses longer than the last of them: at one millisecond the impedance is 0.4090 K/W — 23.7 times less — so a hundred watts for a millisecond raises the junction by 40.9 K rather than by 970. That figure is computed at one millisecond rather than read off the curve, whose nearest reported pulse length is 1.25 ms at 0.488 K/W: sixty lengths spread logarithmically sit 28.8 per cent apart, so a value taken off the grid at a round number is the length after it. The dashed curve is the marched netlist and the solid one the per-stage sum; they part by 0.94 per cent where the stages overlap, which is what says a Cauer ladder's step response is not a sum of its own stages.
Fig. 1 The picture this starts from: the single-pulse thermal impedance of a die, its case and its heatsink, drawn as the marched netlist and as the sum of the stages’ own exponentials. The three shoulders are the three stages filling. It reports the two as 0.94 per cent apart, and it can now say which of them is nearer the truth only because of what is measured below — before that, two curves parting company named no culprit.

That is a true sentence about the arithmetic and it was not a measurement. Two curves that part company tell nobody which of them left. The marched one was called the answer because marching is the more elaborate computation, which is a reason to trust it and not a reason to believe it, and the difference between them was described rather than priced.

The three rungs below all rest on this. The heat a recovery leaves behind computes the energy one switching event costs and turns it into watts; two loops and one heatsink puts two of those on a shared case and iterates to a fixed point. Every one of those numbers is a temperature read off a thermal model, and the model has now been drawn two ways with the difference between them unmeasured.

There is a third route and it is exact.

The ladder's step response, and the sum of its own stages — 0.95 per cent apart at worst. computed by solving, not by drawing. A step of power into a three-stage thermal ladder, and the junction's rise divided by it. The solid curve is exact: the impedance is a continued fraction in s, its denominator has 3 real negative roots, and the partial-fraction expansion of Z(s)/s is a sum of that many ordinary exponentials — no march, no step size. The dashed curve is the sum every account of a thermal path writes, each stage's own resistance times 1 − exp(−t/RC) with its own local time constant, and it is an approximation because the stages load each other. What that costs is 0.950 per cent, once, at 12.9 ms — between the fastest stage's 2.4 ms and the next one's 200 ms, which is the only place two stages are moving together. It is one-sided: the sum never reads low.
Fig. 2 A step of power into the ladder, and the junction’s rise divided by it. The solid curve is the exact step response and the dashed one the sum of the stages’ own exponentials. They part company by 0.950 per cent, once, at 12.9 milliseconds — between the die’s 2.4 and the case’s 200 — and the sum is high there rather than low.

What is being solved

The impedance looking into an RCRC ladder is a continued fraction. The heatsink’s resistance sits in parallel with its own heat capacity; the case’s resistance is in series with that and in parallel with the case’s heat capacity; the die’s resistance is in series with the result and in parallel with the die’s. Three stages give a ratio of two polynomials in ss, the denominator cubic, and building it needs nothing but the six element values and the arithmetic of polynomials.

An RCRC ladder’s poles are real, negative and simple — a network of positive resistances and positive capacitances cannot oscillate and cannot have a repeated natural frequency without being built to — so the partial-fraction expansion of Z(s)/sZ(s)/s is a sum of ordinary exponentials:

z(t)=iRi(1et/τi),τi=1/pi,Ri=N(pi)piD(pi)z(t) = \sum_i R_i\left(1 - e^{-t/\tau_i}\right), \quad \tau_i = -1/p_i, \quad R_i = \frac{-N(p_i)}{p_i D'(p_i)}

No march, no step size, no error that depends on how finely time was cut. And the RiR_i and τi\tau_i that come out are not a formula: they are a second network, a chain of parallel resistor–capacitor sections in series, with the same impedance at its terminals as the ladder it came from. The two objects are the standard pair — a ladder whose nodes are the die, the case and the sink, and a chain whose sections are the three natural frequencies.

The calibration, which comes before anything is claimed

An expansion that finds roots and divides by derivatives is exactly the machinery that returns plausible numbers for a wrong problem, so it is checked three ways before it is asked anything.

The residues must sum to the ladder’s total thermal resistance, because the exponentials have all died by then and z()z(\infty) is that resistance. They sum to 9.7 kelvin per watt, exactly, at the last digit the arithmetic carries — an identity the root-finding could easily have broken and does not.

The poles must be the network’s own natural frequencies, and this collection already computes those another way: it assembles the netlist’s matrices and samples the determinant of G+sCG + sC, a computation that shares no line of arithmetic with a continued fraction. The two agree to 4.5 parts in 101410^{14}. That is the pairing one step computed twice insists on for a transient, applied to the poles rather than to the waveform.

And the third: the ladder’s own state equations, three of them, integrated by taking the matrix exponential rather than by stepping. The junction temperature that comes out of that agrees with the expansion to 1.1×1091.1\times10^{-9} kelvin per watt over nine decades of time. Two routes that share the six element values and nothing else after them — the same standard where the behaviour is written down sets for a netlist and a differential equation describing one circuit.

The last of those three is worth a sentence about why it was done that way rather than by marching. The fastest stage is 2.4 milliseconds and the slowest is 480 seconds, a stiffness of two hundred thousand, and any fixed step over that range must either resolve the first or step over it. Taking the exponential of the system matrix has no step in it at all: it costs about twenty multiplications of a three-by-three matrix and is limited by the arithmetic rather than by a discretisation, which is the same reason a sum that is exact prefers a closed form to a fine grid wherever one exists.

The departure, and it is smaller than it was said to be

With that in place the approximation can be priced instead of characterised.

The worst it does is 0.950 per cent, at 12.9 milliseconds, and the sign of it never changes. Over nine decades of pulse length the sum is high or equal and never low, which is the more useful half of the answer: a closed form that can only over-state a junction temperature is safe to design against, and that is a stronger statement than any bound on its size.

The error is one hump per pair of neighbouring stages, and it is never negative. computed by solving, not by drawing. The local-time-constant sum's error against the exact expansion, as a percentage of the exact answer, over nine decades of pulse length. It is zero at both ends — at times short against every stage only the fastest is charging and there is nothing to load it, and at long times both routes reach the same 9.7 K/W — and it has exactly 2 humps in between, one for each adjacent pair. The tallest is 0.950 per cent at 13 ms; the other is 0.454 per cent at 95 s. Every value is positive, which is the useful half: an approximation that can only over-state a junction temperature is a safe one to design against, and that is a stronger statement than a bound on its size.
Fig. 3 The same disagreement as a curve. It is zero at both ends and has exactly two humps between them — 0.950 per cent at 13 milliseconds and 0.454 per cent at 95 seconds — one for each adjacent pair of stages, each sitting between the two time constants whose overlap causes it.

The shape says what the mechanism is. The error is not a property of the ladder; it is a property of each pair of neighbours, and a three-stage path has two pairs and therefore two humps. Where only one stage is charging there is nothing to load it and the sum is exact; where two are moving together the sum has ignored that one of them is charging into a node that is itself rising.

Fifty per cent, divided by the separation

That mechanism has a size and it depends on one number: how far apart in time the two neighbours are.

What the sum costs is set by one ratio: how far apart the two fastest stages are. computed by solving, not by drawing. The worst error of the local-time-constant sum, against the separation between the fastest stage's time constant and the next one's, swept by changing the die's own heat capacity and nothing else. The straight line is exactly inverse in the separation, anchored at the widest point measured. The measured errors approach it from below: each halving of the separation multiplies the error by 1.94 at the wide end and 1.76 at the narrow one, climbing towards two. So the approximation is not good or bad in itself — it is worth about fifty per cent divided by the separation, which for a die, a case and a heatsink three orders apart is under a per cent and for two stages a factor of five apart is 10.7.
Fig. 4 The worst error against the separation of the two fastest time constants, swept by changing the die’s own heat capacity and nothing else. Each halving of the separation multiplies the error by 1.94 at the wide end and 1.76 at the narrow one, climbing towards two — so the error is asymptotically inverse in the separation, with a coefficient near fifty per cent.

A die at 2.4 milliseconds, a case at 200 and a heatsink at 480 seconds are 83 and 2,400 times apart, and 50 per cent divided by 83 is the 0.95 measured above. A path whose two fastest stages sit a factor of five apart gives 10.7 per cent, and there the closed form genuinely is the wrong object.

So the honest statement about the approximation is neither that it is good nor that it is an approximation. It is that it costs about half divided by the separation, and that a real package — a die of milligrams, a case of grams, a heatsink of kilograms — separates its stages by construction. The sum is accurate for the same reason the thermal model is a ladder in the first place, which is a better sort of reason than an empirical one.

Why the separations are large, and what would make them small

The coefficient is a fact about arithmetic; the separations are a fact about parts, and they are large for a reason that has nothing to do with the electrical problem.

A thermal time constant is a resistance times a heat capacity, and a heat capacity is a mass. A die weighs milligrams, its package and the grease under it weigh grams, and a heatsink weighs hundreds. Each stage therefore carries about three orders of magnitude more heat than the one above it, while the resistances move the other way by rather less, and what is left is the two factors of 83 and 2,400 that the ladder has. A thermal path is stiff by construction, and the closed form is accurate for the same reason the lumped model is legitimate at all.

The case where it is not is where two stages are built to be similar — a die attached directly to a copper slug of comparable mass, or a two-layer model of a single die’s interior, which is what a hot-spot calculation needs. There the separation falls to a factor of a few, the error goes to the ten per cent measured above, and the closed form has stopped being a convenience. The diode that conducts backwards is the event that would need such a model: a recovery is tens of nanoseconds, and at that scale the object being heated is a few square microns of silicon and not a die at all.

The second network is not the first one rearranged

The expansion produces element values, and it is worth looking at what they are.

Two networks, one impedance: every element differs and the terminal is identical. computed by solving, not by drawing. The Foster chain that reproduces this ladder exactly, compared element by element with the ladder it came from. Each stage contributes two bars — its section's time constant against the stage's own RC, and its section's resistance against the stage's own R — and not one of the six is zero. The die's section is -0.503 per cent off in time constant and -1.010 per cent off in resistance, and the differences go both ways rather than one. The two networks nevertheless have the same impedance at their terminals to every digit the arithmetic holds, and the resistances still sum to 9.7 K/W exactly, because the expansion is an identity rather than a fit. What has been lost is the correspondence: a Foster section is a pole, not a piece of the package.
Fig. 5 Every element of the chain against the ladder stage it stands for. The die’s section is 0.503 per cent short in time constant and 1.010 per cent short in resistance; the case’s is 0.163 short and 1.082 over; the heatsink’s 0.670 over and 0.084 over. The resistances still sum to 9.7 kelvin per watt exactly.

Six element values and not one of them survives. The differences are small — a per cent — and they go both ways, so they are not a scaling that could be undone. The chain is 1.18787 kelvin per watt and 2.38792 milliseconds where the ladder is 1.2 and 2.4; 0.505409 and 199.674 milliseconds where the ladder is 0.5 and 200; 8.00672 and 483.217 seconds where the ladder is 8 and 480.

And yet the two networks are indistinguishable at their terminals. That is not an approximation holding well; it is an identity. The impedance is one rational function and the two networks are two factorisations of it, in the same way that the two numbers a source can be described by are two factorisations of one linear behaviour.

A pole is not a place

Which raises the question the pair of networks exists to answer, and it has a sharper answer than “the element values differ”.

A data sheet does not publish a die’s heat capacity. It publishes a transient impedance curve, and what can be extracted from a curve is exactly the chain: a handful of resistance–time-constant pairs that reproduce it. Those pairs are then used in a simulator, and the simulator has nodes in it, and the node between the first section and the second looks like the case.

It is not the case.

The Foster chain's inner node is not the case: one rises as t, the other as t². computed by solving, not by drawing. The case temperature of the ladder, solved from its own state equations, against the potential at the corresponding point inside the Foster chain — the node after the first section, which is what anyone reading a data sheet's Foster pairs would take the case to be. They end at the same place and they start differently in KIND. The real case is two integrations from the power that heats it and rises as time to the 1.996; the Foster node is one section from the terminal and rises as time to the 1.000. At 100 µs the difference is 49.6 times. A Foster network is an exact model of the junction and has no interior at all: its sections are poles, and a pole is not a place.
Fig. 6 The ladder’s real case temperature against the potential at the corresponding point inside the chain. They end at the same 8.5 kelvin per watt and they begin differently in kind: the real case rises as time to the 1.996 and the chain’s node as time to the 1.000, both fitted over a decade with a coefficient of determination of unity to six places. At 100 microseconds the chain’s node reads 49.6 times the real rise.

The exponents are the whole content. A real case is two integrations away from the power that heats it — the power charges the die, the die’s excess drives current through the first resistance, that current charges the case — so its rise starts as t2t^2. The sink is three integrations away and starts as t3t^3; fitted over the same window it comes out at 2.984. The chain has no such structure: every one of its sections hangs directly off the terminal, so every internal node is one integration from the source and starts as tt.

The consequences are quantitative and they are large where it matters. The chain’s inner node reads 5.59 times the real case rise at a millisecond, 1.33 times at ten, and comes within ten per cent only past 32.3 milliseconds and within one per cent only past 45.2 seconds. A converter switching at any ordinary frequency lives entirely inside the region where the number is wrong by an order.

This is the same defect exact outside and wrong within collects, arriving from an unusual direction: the model is not approximate and then wrong beyond its range — it is exact for the quantity it was fitted to and meaningless for every other quantity in it. Nothing in the curve it came from ever contained a case temperature, so nothing in the network built from that curve can.

The largest disagreement is not about heat

There is a third source of difference in a picture of this kind and it turns out to be the biggest one, and it is not a model at all.

The march is right to a part in 21169 and reading it off costs 2.9 per cent. computed by solving, not by drawing. Three errors against the exact expansion, all of the same marched run. The lowest curve is the march at the instants it computed: 4.72e-3 per cent at worst, over 848 samples, which is the trapezoidal rule doing its job. The top curve is that run reported at sixty logarithmically spaced pulse lengths by taking the first sample at or after each — 2.919 per cent at 213 µs, because the march steps uniformly inside a decade while the report is logarithmic, so near the bottom of a decade the nearest later sample is a whole step late. The middle curve interpolates between the two samples either side and costs 0.0152 per cent. The largest disagreement in a picture of this kind is therefore not between two models of the heat flow at all.
Fig. 7 Three errors against the exact answer, all of one marched run. At the instants it computed, the march is right to 4.72 thousandths of a per cent — a part in 21,169. Reported at sixty logarithmically spaced pulse lengths by taking the first sample at or after each, the same run is 2.919 per cent wrong at 213 microseconds. Interpolating between the two samples either side costs 0.0152.

The march is essentially exact. It steps uniformly inside each decade, hands its state on at the boundary, and its worst sample over 848 of them is out by five parts in a hundred thousand — the trapezoidal rule doing what it is supposed to and no more. What costs three per cent is asking that run for a value it did not compute and being given the next one it did. Near the bottom of a decade the next sample is a whole uniform step later, and on a curve that is climbing steeply that is three per cent of impedance.

The same arithmetic applies to the reported grid itself. Sixty rows over six and a half decades sit 28.8 per cent apart in time, so a value read off such a curve at “one millisecond” is the row at 1.2526 milliseconds. The exact impedance at a millisecond is 0.40897 kelvin per watt and at that row it is 0.488056 — 19.3 per cent higher — before any read-back error is added.

It applies to the departure above as well, which is why the two pictures on this page disagree about it. 0.950 per cent at 12.91 milliseconds is what the figure at the head of this essay measures: the exact expansion against the sum, scanned at fifteen hundred pulse lengths. The sixty-length picture reports 0.94, and the whole of the difference is grid and route — its lengths straddle the peak, the nearest sitting at 12.19 milliseconds where the departure is 0.9495, and the curve it subtracts is the march rather than the expansion, which sits a part in six thousand towards the sum there and closes the gap a little further. Two grids and a substituted route, for 1.2 per cent of the answer. The figure above is the one this essay means throughout, because it is the only one of the two with an exact curve in it.

None of that is heat flow. It is a logarithmic axis read off a linear grid, and it was three times the closed form’s error and six hundred times the march’s. Ranking the three was the point: 2.919 per cent from the reading, 0.950 from the model, 0.0047 from the integration. A picture in which two curves visibly separate invites the reader to attribute the gap to the two models it labels, and most of that one belonged to neither.

Past tense, because the middle curve is the repair and it is one line. Interpolating between the two samples either side of each reported time takes the worst reported error from 2.919 per cent to 0.0152, at no cost in anything — the samples are dense enough that a step is under two per cent of the elapsed time, so a straight line across one of them is wrong by less than the march is. The single-pulse curve at the head of this essay is the interpolated one; it was the nearest-later one until this measurement was made, and the two-route comparison drawn on it was reporting a defect in the reading as a difference between two models of heat flow.

What it does not say

It does not say the sum of exponentials should be replaced. Inside a package whose stages are two or three orders apart it is under one per cent and always conservative, it costs three exponentials to evaluate, and it superposes — which is the property a pulse train needs and the reason data sheets publish it in the first place. A duty cycle’s worth of steps added and subtracted is a sum of terms in the same three time constants, where the same calculation on a ladder is another march.

It does not say the march was wrong. It was right to five parts in a hundred thousand, which is far better than anything the model’s own linearity deserves — the sink’s convective resistance falls with temperature and the die’s heat capacity rises with it, corrections the rung below puts at of order ten per cent over the range these numbers cover, which is the edge that is a region rather than a line.

And it does not say the chain is a bad model. It is an exact one, of the junction. What it has no opinion about is the inside, and the failure is in reading a network as a picture of a package rather than as a factorisation of an impedance.

Nor does the ranking of the three errors carry over unchanged to a different question. It is a ranking of the three at their worst, and the three have their worst in different places: the model’s is between two time constants, the reading’s is at the bottom of a decade, and the march’s is wherever the curvature is greatest. At a pulse of one millisecond the closed form is 0.599 per cent high and the reading 0.389; at ten seconds they are 0.413 and 0.242. The order of the two reverses away from the reading’s own worst place, so a single number for “the accuracy of the curve” is the wrong shape of answer, in the way every model has an edge is about.

What it opens

Two things, and both are about identification rather than about heat.

The first is that the direction of the derivation only runs one way. Given a ladder, the chain follows in closed form. Given a chain — which is all a measured curve can yield — the ladder does not follow, because the three sections carry three resistances and three time constants and the ladder has six free values with a different meaning. Recovering a physical ladder needs information the terminal never had, which is the same shortfall two numbers without solving for the waveform runs into from the other side.

The second is that the pricing above is a template. Any place this collection compares a closed form with a march is a comparison of two inexact things, and the ranking found here — reading, then model, then integration — is not obviously peculiar to a thermal path. It is the ordering that follows whenever the reported grid is coarser than the computed one, which is nearly always, and it is invisible without a third route.

The number worth carrying

Half a per cent divided by the separation, in a picture whose read-back cost three until it was interpolated.

The habit that goes with it is the harder half. Two computations of one quantity that disagree do not say which is wrong, and the older, cheaper, more approximate-looking one is not reliably the guilty party — here it was right to a per cent while the reading of the elaborate one was wrong by three, and the repair belonged to neither of them. What settles it is an exact case, and for a network of resistances and capacitances an exact case is usually available: the poles are real, the expansion is finite, and the identity that the residues sum to the direct-current answer is a check the arithmetic cannot pass by accident.

The second habit is narrower and is worth stating because it cost three per cent here. A computed curve and a reported curve are different objects, and the grid a result is printed on is part of the result. Sixty rows over six and a half decades cannot resolve anything narrower than 29 per cent in time, however exactly each row was computed, and a number lifted off such a grid at a round value carries that spacing with it — the digits the arithmetic did not have, arriving one level up, where what is missing is not precision but resolution.

Part 5 on Reverse-recovery

One argument about Reverse-recovery, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

Closed formContinued fraction expansionLumped approximationMarchingModel rangePolesResiduesThermal impedanceThermal resistanceVerification