We read the footnotesIssue 01 · October 2026
The Nerd Uncle
Issue 01 › Economics

The clock you are actually paying for.

Latency looks like difficulty and almost never is. Separating the two changes which model you should be buying, and how much of it.

We read the footnotes · 4 min read

Horizontal bars, each mostly pale grey with a short red segment at the end.
Each row is one request. The pale span is time queued before generation begins; the solid span is generation itself. The ratio is the finding.

Two clocks, not one

There is the gap before the first word appears, and there is the rate at which the rest arrives. Most teams quote a single latency figure that averages the two together, and then spend months optimising whichever one happened to be innocent.

Two spans, measured apart. Averaging them into one number is how the wrong half gets optimised.
Two spans, measured apart. Averaging them into one number is how the wrong half gets optimised.

The first gap is mostly scheduling

Before a single word is generated, your request joins a queue shared with everyone else pointed at that endpoint. At busy hours that scheduling delay dominates everything else, and no amount of prompt tuning will shorten it.

This is also why a benchmark run at nine in the morning tells you almost nothing about four in the afternoon.

The rest is for subscribers.

It carries on in the same plain language, with the same drawn diagrams.

One week free, then ₹25