Time to First Token (TTFT)
Also known as: TTFT
Latency from request to the first token landing in the client; the dominant chat UX latency metric.
Latency from request submission to the first generated token reaching the client. Driven by prefill compute (proportional to prompt length) plus queue wait. The dominant perceived latency metric for chat apps.
In practice
First number any production LLM team tracks. Interviews probe prefill vs decode, batching impact, and prompt-caching gains.
Related topics
Questions that mention this term
Related terms
API LLM
An LLM accessed through a provider API: pay per token, get the frontier model, hand over ops.
Beam Search
Keep the K best partial sequences at each step; deterministic, breadth-first decoding.
FlashAttention
A memory-aware attention kernel that's 2-4x faster than vanilla, with identical math.
GGUF
Self-contained binary format for quantized LLMs; the standard for llama.cpp / Ollama / LM Studio.
Greedy Decoding
At each step, pick the single highest-probability token. Fast and deterministic, but often loops.
Grouped-Query Attention (GQA)
Compromise between MHA and MQA: query heads share KV heads in groups, cutting KV cache by 4-8x.