Top-p (Nucleus) Sampling
Also known as: Nucleus sampling
Sample from the smallest token set whose probabilities sum to ≥p; an adaptive alternative to top-k.
A decoding strategy that samples the next token from the smallest set whose cumulative probability exceeds p. Adapts the candidate pool size dynamically: broad when many tokens are plausible, narrow when one dominates.
In practice
One of the two main decoding knobs (with temperature). Interviews probe its interaction with temperature and why low-p + temp=0 are duplicative.
How it compares
Temperature reshapes the probability curve; top-p truncates the tail before sampling.
Related topics
Questions that mention this term
Related terms
API LLM
An LLM accessed through a provider API: pay per token, get the frontier model, hand over ops.
Beam Search
Keep the K best partial sequences at each step; deterministic, breadth-first decoding.
FlashAttention
A memory-aware attention kernel that's 2-4x faster than vanilla, with identical math.
GGUF
Self-contained binary format for quantized LLMs; the standard for llama.cpp / Ollama / LM Studio.
Greedy Decoding
At each step, pick the single highest-probability token. Fast and deterministic, but often loops.
Grouped-Query Attention (GQA)
Compromise between MHA and MQA: query heads share KV heads in groups, cutting KV cache by 4-8x.