The clock you are actually paying for.
Latency looks like difficulty and almost never is. Separating the two changes which model you should be buying, and how much of it.

Two clocks, not one
There is the gap before the first word appears, and there is the rate at which the rest arrives. Most teams quote a single latency figure that averages the two together, and then spend months optimising whichever one happened to be innocent.

The first gap is mostly scheduling
Before a single word is generated, your request joins a queue shared with everyone else pointed at that endpoint. At busy hours that scheduling delay dominates everything else, and no amount of prompt tuning will shorten it.
This is also why a benchmark run at nine in the morning tells you almost nothing about four in the afternoon.
The rest is for subscribers.
It carries on in the same plain language, with the same drawn diagrams.
One week free, then ₹25