latency

TTFT

Time to first token. The pause between your question and the model’s first word. The only latency number a human actually feels.

Every inference system has two clocks. Throughput tells you how many tokens per second the hardware can push. TTFT tells you how long someone waits before they see a single word appear.

When TTFT is low, a model feels alive — like it’s thinking with you. When it’s high, even a brilliant answer feels broken. Prefill, scheduling, queueing, network, the first hop of the decode loop: it all shows up in this one number.

This domain is about treating that pause as a first-class engineering problem. Because at scale, the difference between 80 ms and 800 ms is the difference between a tool people use and a tool people tolerate.

tpot.siTime per output token — the clock after the first word. kvcache.siThe memory that decides how long that first word takes.