Circuits that do a job, and the range they do it over

The decision taken where the ramp is slowest

A relaxation oscillator decides at its thresholds, and a threshold is the one place on a charging exponential where the slope is smallest. Noise there costs 3.399 units of period against the 2.828 that counting two crossings gives, because a draw moves the crossing it is armed for and the level the next ramp starts from. Consecutive periods share that draw, so they are positively correlated and the jitter accumulates at 3.771 per root period rather than at 3.399. And at a fixed frequency there is a best hysteresis: β = 0.648, where β·ln((1+β)/(1−β)) = 1.

Assumes: Two thresholds because there is a floor · The gain that is exactly one

Two thresholds because there is a floor put hysteresis on a comparator to stop it changing its mind twenty-two times on one crossing. That essay was about a decision that has to be made once. The period a delay lengthens fed the same comparator’s output back through a resistor and turned the pair of thresholds into a period, and the delay that is two delays made the two propagation delays different and measured what that does to the duty cycle.

None of the three has noise in it once the oscillator is running. The first essay’s noise is there to be defeated; the second and third are exact marches of a noiseless circuit. But the noise did not go away when the hysteresis was added, and in an oscillator the decision is not made once — it is made twice a period, for ever, and every one of those decisions is taken at the point on the capacitor’s charging curve where the slope is smallest.

The threshold is the worst place on the ramp to ask a question, and the cost is measurable. At a divider ratio of a half the period’s standard deviation is 3.399 units of RC·σ/V, against the 2.828 that the obvious estimate gives — and the long-term jitter accumulates at 3.771 per root period, which is above both.

The period's spread is 1.20 times what counting two crossings givescomputed by solving, not by drawing. 19999 periods of a relaxation oscillator with 5 mV rms of noise on its thresholds, computed from the exact flip instants rather than marched, at β = 0.5. The measured standard deviation is 3.4 ns and the closed form — three partial derivatives of the period with respect to the three draws it depends on — gives 3.4 ns. The estimate that counts two threshold crossings and divides the noise by the slope at each gives 2.83 ns, which is 17 per cent low. The curve is the closed form's Gaussian, drawn on the measured histogram rather than fitted to it.05e+71e+8-10010period, in nanoseconds either side of the noiseless 2.2 µsdensity (per second)the closed form, not fitteddivider ratio β0.5threshold noise5 mV rmsperiods measured19999measured spread3.4 nsclosed form3.4 nsthe two-crossing estimate2.83 nsnoiseless period2.2 µssolved, then checked — an event map against three derivativesthey agree to 0.1%
Fig. 1 Twenty thousand periods of the oscillator with 5 mV rms of noise on its thresholds, computed from exact flip instants rather than marched, at β = 0.5. The measured spread is 3.40 ns; three partial derivatives give 3.40; the estimate that counts two crossings gives 2.83. The curve is the closed form’s Gaussian drawn on the histogram rather than fitted to it. The slider is the noise.

Why the march had to be put down

The three earlier essays walk the network forward with the trapezoidal rule, and that is the right instrument for a waveform. It is the wrong one for a statistic.

A standard deviation to three figures needs tens of thousands of periods, and a march is a couple of thousand steps in each of them. What makes that avoidable is that nothing happens between two flips that the network does not already say in closed form: the capacitor charges towards one rail from wherever the last flip left it, so each half cycle is one logarithm. The flip instants are exact and the record is as long as it needs to be.

The model of the noise is the one the march already had, so the two remain comparable: the threshold the comparator is waiting for is β·V less a fresh normal draw, armed when the threshold is armed and held until it is crossed. That is noise with a correlation time long against the crossing and short against a half cycle — a comparator’s own input offset drifting on the timescale of the oscillation, rather than the wideband floor the floor a circuit has measures. It is the one assumption in this essay that a reader should hold loosely, and the last section says what a different assumption would change.

A draw does two things, and the estimate counts one

The estimate everybody makes is the obvious one. Noise on a threshold turns into an uncertainty in when the threshold is crossed, at a rate set by the slope there. Two crossings a period, independent, so the period’s standard deviation is 2\sqrt{2} times the per-crossing figure.

The slope at a threshold is V(1 − β)/RC, which is the slowest the ramp ever goes — the capacitor is heading for the rail and the threshold is the closest it comes to it. So the per-crossing timing error is σ·RC/(V(1 − β)), and the estimate is 2\sqrt{2} times that: 2.828 units at β = 0.5.

What the estimate misses is that the comparator does not merely decide late or early. It leaves the capacitor somewhere. A draw of n makes the flip happen at β·V − n rather than at β·V, so the next half cycle starts from the wrong level, and it is longer by n divided by the slope the ramp starts with — which is V(1 + β)/RC, the fastest it ever goes.

Writing A = RC/(V(1 + β)) for that second sensitivity and B = RC/(V(1 − β)) for the first, a period is two half cycles and therefore three draws, with coefficients A, A + B and B. The middle one is the draw shared between the two halves and it is the largest of the three. The period’s standard deviation is σA2+(A+B)2+B2\sigma\sqrt{A^2 + (A+B)^2 + B^2}, which at β = 0.5 is 3.399 units — twenty per cent more than the estimate, and the whole of that twenty per cent is the term about where the next ramp starts.

Measured over twenty thousand periods of the event map, with nothing shared between the two routes but the value of the noise, the standard deviation is 3.40 nanoseconds against the closed form’s 3.40.

Counting two crossings gets 17 per cent of the jitter at β = 0.5, and less at every other ratio. computed by solving, not by drawing. Three quantities against the divider ratio, all in units of RC·σ/V so that β alone decides them. The lowest curve is the estimate that counts two threshold crossings and divides the noise by the slope at each; the middle one is the true standard deviation of the period, which also carries the draw that decides where the NEXT ramp starts; the highest is the rate at which jitter accumulates over many periods, which is larger still because consecutive periods are positively correlated. At β = 0.5 they are 2.8284, 3.3993 and 3.7712. All three diverge as β approaches one, where the ramp arrives at its threshold almost flat.
Fig. 2 Three quantities against the divider ratio, in units of RC·σ/V so that β alone decides them: the two-crossing estimate, the period’s true standard deviation, and the rate at which jitter accumulates. At β = 0.5 they are 2.8284, 3.3993 and 3.7712. All three diverge as β approaches one, where the ramp arrives at its threshold almost flat.

More hysteresis is not more jitter, and it is not less either

Read along the axis of that figure and the shape is not the one the first essay’s argument suggests.

More hysteresis means a threshold further from the rail, so the ramp arrives at it more slowly, so the same noise is worth more time: B = RC/(V(1−β)) grows without bound as β approaches one. At β = 0.8 the period’s spread is 7.49 units against 3.40 at a half, and at β = 0.9 it is worse again.

That looks like an argument for small hysteresis and it is not one, because the period shrinks as well. An oscillator with β = 0.1 has a period of 0.2 RC where one with β = 0.8 has 4.4 RC, so the same jitter is a much larger fraction of a much shorter cycle. Comparing two oscillators means comparing them at something, and the honest comparison is at a fixed frequency — which the last figure in this essay does, and which turns the whole argument around.

What can be said without fixing anything is the ratio between the three curves. The two-crossing estimate is low by twenty per cent at β = 0.5, by thirty-one per cent at β = 0.2 and by six per cent at β = 0.8: the term it forgets matters most where the hysteresis is small, because there A and B are comparable and the forgotten sensitivity is nearly as large as the counted one.

The period's spread is 1.06 times what counting two crossings gives. computed by solving, not by drawing. 19999 periods of a relaxation oscillator with 5 mV rms of noise on its thresholds, computed from the exact flip instants rather than marched, at β = 0.8. The measured standard deviation is 7.49 ns and the closed form — three partial derivatives of the period with respect to the three draws it depends on — gives 7.49 ns. The estimate that counts two threshold crossings and divides the noise by the slope at each gives 7.07 ns, which is 6 per cent low. The curve is the closed form's Gaussian, drawn on the measured histogram rather than fitted to it.
Fig. 3 The same measurement at β = 0.8. The period’s spread is 7.49 ns against the two-crossing estimate’s 7.07 — six per cent, against twenty at β = 0.5 — because with a threshold close to the rail the crossing’s own sensitivity swamps the one that decides where the next ramp starts.

The shared draw, and what it does to a clock

The draw in the middle of a period is the one that ends the first half cycle and starts the second. It is also the draw that ends the previous period and starts this one, which means consecutive periods are not independent.

They are positively correlated. The covariance is ABAB and the variance is A2+(A+B)2+B2A^2 + (A+B)^2 + B^2, so the correlation is AB/(A2+(A+B)2+B2)AB/(A^2 + (A+B)^2 + B^2)0.115 at β = 0.5, and the event map measures 0.117 over twenty thousand periods. It is positive at every divider ratio, largest where the hysteresis is smallest, and it approaches a sixth as β vanishes and the two sensitivities become equal. At the other end it falls to nothing, because a threshold near the rail makes B swamp A and the shared draw stops mattering to the half cycle it starts.

The consequence is the one a clock designer cares about. A long-term jitter is usually got from a per-period measurement by multiplying by the square root of the number of periods, and that step assumes independence. With a positive correlation the errors partly add instead of partly cancelling, and the accumulated jitter grows faster than the per-period figure predicts.

How much faster is exact. Every interior draw appears in two consecutive half cycles with total coefficient A + B, and there are two of them a period, so the standard deviation of the Nth edge’s position is σ2N(A+B)\sigma\sqrt{2N}\,(A + B)3.771 units per root period at β = 0.5, against the 3.399 a naive N\sqrt{N} would give. Ten per cent, every time, for ever.

One period and the next share a decision, and a sixth is the most that can buy. computed by solving, not by drawing. The correlation between consecutive periods against the divider ratio, from the three partial derivatives. It is AB/(A² + (A+B)² + B²), where A and B are how much a draw moves the ramp it starts and the crossing it ends; it is positive everywhere, peaking at 0.164 at the smallest hysteresis drawn — approaching a sixth as β vanishes and the two sensitivities become equal — and falling to nothing as β approaches one, where the crossing's own sensitivity swamps the other. A positive correlation is why a long-term jitter cannot be got by multiplying a per-period figure by the square root of the count: the errors partly add rather than partly cancelling.
Fig. 4 The correlation between one period and the next against the divider ratio. Positive everywhere, approaching a sixth as the hysteresis vanishes and the two sensitivities become equal, and falling to nothing as β approaches one. It is why an accumulated jitter cannot be got from a per-period figure by multiplying by N\sqrt{N}.

Measured over a hundred and twenty runs

The accumulation is the claim worth checking against something rather than deriving twice, because the derivation is where a sign error would hide.

A hundred and twenty independent seeded runs of the event map, each giving the position of the Nth rising edge relative to the first, produce a spread that grows as the square root of N and sits on the predicted line to six per cent — which is about what a hundred and twenty samples buys. At a hundred periods the measured spread is 36.6 nanoseconds, the shared-draw prediction is 37.7, and the per-period figure times N\sqrt{N} is 34.0.

Three nanoseconds in a hundred periods is not, on its own, a reason to change anything. What the figure is for is the shape: the error from using N\sqrt{N} does not go away with more periods and does not average out, because it is a wrong coefficient on a correct square root. A jitter budget written over a thousand periods is low by the same ten per cent as one written over ten.

It also says what a per-period measurement is and is not evidence for. An instrument that reports cycle-to-cycle jitter is measuring the 3.399; a system that cares when an edge arrives after a thousand cycles is exposed to the 3.771. The two numbers are about the same oscillator and neither is the other’s approximation.

The jitter accumulates faster than the per-period figure times √N. computed by solving, not by drawing. The standard deviation of the position of the Nth rising edge, over 120 independent seeded runs of the event map, against N — for β = 0.5 and 5 mV of threshold noise. The points are measured; the solid line is √(2N) times the coefficient of the draw that ends one half cycle and starts the next; the dashed line is the per-period standard deviation times √N, which is what a measurement of one period would suggest and which is low by 10 per cent. Measured and predicted agree to 6 per cent, which is what 120 runs buys.
Fig. 5 The spread of the Nth rising edge’s position, over 120 seeded runs, against N. The solid line is 2N\sqrt{2N} times the shared draw’s coefficient; the dashed line is the per-period spread times N\sqrt{N}, which is low by ten per cent at every count. Measured and predicted agree to six per cent.

The linearisation, and the refusal that ends it

Three partial derivatives are a linearisation, and a linearisation needs a statement of where it stops.

Swept from half a millivolt to five hundred — a thousandfold, and from two hundredths of a per cent of the threshold voltage to twenty per cent — the measured spread and the closed form agree to 4.8 per cent at the worst point and to a tenth of one per cent over most of the range. At a fifth of the threshold the form is still within five per cent, which is further than it has any right to hold, and the departure is the second-order term arriving rather than anything qualitative.

What ends the model is not a drift. At about a third of the threshold a draw is eventually large enough to put the armed threshold outside the rail the capacitor is charging towards, and then the capacitor never reaches it: the oscillator stops. The event map refuses that case by name instead of returning a number, which is the correct behaviour for a solver whose answer would otherwise be a plausible period for an oscillator that has latched.

That failure is not an artefact of the arithmetic. It is the same failure two thresholds because there is a floor is about, arriving from the other side: hysteresis large against the noise defeats chatter, and noise large against the hysteresis defeats the oscillator. Between them is the whole design, and this essay’s sweep says the useful part of it extends a great deal further into the noisy end than a designer would guess.

The linearisation holds to 5 per cent over a thousandfold of noise, and then the oscillator stops. computed by solving, not by drawing. The standard deviation of the period against the size of the threshold noise, from 12,000 half cycles of the event map at each level (points), against the closed form (dashed), for β = 0.5. The two agree to 4.8 per cent across the whole sweep, which covers noise from 0.02 to 20 per cent of the threshold voltage. The closed form is a linearisation and it does not stop being right where a reader expects it to: it is still within five per cent when the noise is a fifth of the threshold. What ends it is not a drift but a refusal — at about a third of the threshold a draw is eventually large enough to put the armed threshold outside the rail the capacitor is charging towards, and the event map says so by name rather than returning a period.
Fig. 6 The period’s spread against the size of the threshold noise, measured on the event map and computed from the derivatives, across a thousandfold. They agree to 4.8 per cent at the worst point. The sweep stops where it does because above about a third of the threshold the oscillator latches, which the solver refuses rather than reporting.

The best hysteresis, at a frequency somebody asked for

A designer does not choose β and then find out what frequency they got. They are given a frequency and choose the parts, and at a fixed frequency the two effects of hysteresis pull against each other.

More hysteresis slows the ramp at the threshold, which costs jitter. More hysteresis also lengthens the period for a given RC, so holding the frequency means a smaller RC, which buys jitter back — the timing error is σ·RC/V times a function of β, and RC is now proportional to 1/ln((1+β)/(1−β)).

Putting the two together, the accumulated jitter per root period at a fixed period T is

2σT/V\sqrt{2}\,\sigma T/V divided by (1β2)ln((1+β)/(1β))(1-\beta^2)\ln((1+\beta)/(1-\beta)),

and the denominator has a maximum. Differentiating it gives 2 − 2β·ln((1+β)/(1−β)), so the optimum is where β·ln((1+β)/(1−β)) = 1, which has exactly one root between zero and one: β = 0.6479.

It is a real minimum rather than a flat region. A tenth is 4.51 times worse; 0.95 is 2.51 times worse; the familiar half is 8.7 per cent worse, and the curve is within ten per cent of its best between about 0.50 and 0.78. So a divider ratio chosen for convenience is not being punished much, and a ratio chosen because it looked like plenty of hysteresis may be.

Two routes reach that number and they share only the expression for the period. Scanning the curve over forty-nine ratios finds its lowest point at 0.66; bisecting the stationary condition finds 0.6479. The scan is coarse and is meant to be: what it is checking is that the condition is the minimum of the thing it was derived from rather than an algebraic slip.

At a fixed frequency there is a best hysteresis, and it is β = 0.648. computed by solving, not by drawing. The rate at which jitter accumulates, per root period, against the divider ratio — with the RC chosen at each ratio to keep the frequency the same, so that the comparison is between oscillators a designer could swap for one another. More hysteresis makes each decision less noisy and at the same time forces a smaller RC to hold the frequency, and the two pull in opposite directions. The minimum is at β = 0.6479, where β·ln((1+β)/(1−β)) = 1, and the value there is 1.5793 σT/V. It is a real minimum: β = 0.1 is 4.51 times worse and β = 0.95 is 2.51 times worse, and the curve is within ten per cent of its best between about 0.50 and 0.78.
Fig. 7 Accumulated jitter per root period against the divider ratio, with the RC chosen at each ratio to hold the frequency. The minimum is at β = 0.648, where β·ln((1+β)/(1−β)) = 1, and the curve is within ten per cent of its best between 0.50 and 0.78.

What the numbers are on a real part

Everything above is in units of RC·σ/V, which makes β the only variable and hides how large any of it is. Putting a part in is worth doing once.

A general-purpose comparator with twenty nanovolts per root hertz of input noise and ten megahertz of bandwidth has about 79 microvolts rms at its input when that density is integrated over its own noise bandwidth — which is the integral the bandwidth noise sees is about. In an oscillator with RC = 1 µs and five-volt rails at β = 0.5, the unit RC·σ/V is 15.85 picoseconds, so the period’s standard deviation is 53.9 picoseconds on a 2.197-microsecond period — twenty-four parts per million, cycle to cycle. The two-crossing estimate would have said 44.8.

Over a thousand periods the accumulated jitter is 1.89 nanoseconds, which is what matters if anything downstream counts cycles. Moving to the optimum β of 0.648 at the same frequency needs RC = 0.712 µs and gives 50.9 picoseconds a period and 55.0 per root period — five and a half per cent better than the half, which is the size of the prize the last figure sizes and is worth having only where it is free.

Two cautions go with those figures. The noise model is the slow one — a comparator’s offset drifting on the timescale of the oscillation — and integrating a wideband density to get its size is a convenience rather than a derivation; a part whose noise is genuinely white at the crossing obeys the other model named below. And the input noise is not the only contributor: the supply’s noise reaches the thresholds through the divider, and the reference the divider divides is usually the same rail the comparator swings.

What this leaves the two delay results

Two earlier results about the same oscillator are affected and neither is overturned.

The delay still lengthens the period, by 4/(1+β) times itself rather than by twice it, and the noise does not interact with that: the delay adds a fixed time to each half cycle and the draws add a random one, and the two sensitivities are computed at different points of the ramp. What the noise does add is a reason to want the delay small that has nothing to do with the frequency error — a comparator with a long delay needs a smaller RC for the same frequency, and a smaller RC is more jitter.

And the duty cycle’s skew is unaffected on average and not in spread. The delay that is two delays measures a duty cycle moved by 0.28 per cent for twenty nanoseconds of skew. The noise puts a spread on the duty cycle as well as on the period, and because a duty cycle is the difference of two half cycles rather than their sum, the shared draw enters it with the opposite sign — which is a distinct calculation from the one here and is the first thing to do next.

Still open: the duty cycle’s own spread, a wideband floor, and the walk into the core

The duty cycle’s spread, where the shared draw subtracts. A period is H++HH_+ + H_- and a duty error is H+HH_+ - H_-, so the draw that both halves share enters the first with coefficient A + B and the second with B − A. At β = 0.5 that is 2.667 against 1.333, so the duty cycle should be quieter than the period by a factor that depends on β and vanishes as β does. Measuring it would say whether the 0.28 per cent of duty error a twenty-nanosecond skew produces is visible above the noise on a real comparator, which is the question that decides whether the skew is worth trimming.

A wideband floor rather than a slow offset. The noise here is one draw per decision, held while the threshold is armed. The opposite idealisation is a white floor whose correlation time is short against the crossing, and there the timing error is not the noise over the slope — it is set by the noise bandwidth against the slope, and the crossing is a level-crossing problem rather than a displaced threshold. The two models differ by a factor that contains the comparator’s bandwidth, and which of them a real part obeys is decided by where its dominant noise is. The measurement that separates them is the dependence on β: the model here gives the three coefficients above, and a bandwidth-limited one should give a different mixture.

The walk into the core, which is where the duty-cycle argument was already pointing. The delay that is two delays found ten nanoseconds of comparator skew saturating a hundred-turn core in 231 cycles, because a fixed volt-second imbalance accumulates, and a boundary in volt-seconds is where that limit is drawn. A random imbalance accumulates too, as a walk rather than a ramp, so the flux reaches a given level in a number of cycles proportional to the square of it rather than linearly. With the two together a core sees a ramp plus a walk, and which of them arrives first depends on the ratio of the skew to the noise — a boundary in nanoseconds per millivolt that neither essay has drawn.

Part 4 on hysteresis

One argument about Hysteresis, and one of 5 essays on it so far, each part numbered by how much of the idea it assumes. What sits either side of it:

The objects named here

The third axis, after the field and the idea: the things themselves, and every essay that touches each one.

ComparatorDesign tradeoffHysteresisNoise floorPeriod jitterRelaxation oscillatorSeeded generatorVerification