Beam Search
Also known as: Beam decoding
Keep the K best partial sequences at each step; deterministic, breadth-first decoding.
A deterministic decoding strategy that maintains the top-K highest-probability partial sequences (beams) at each step. Trades exploration for breadth; often produces bland or repetitive output for open-ended generation.
In practice
Standard for translation/summarization, rarely for chat. Interviews probe why temperature sampling won for open-ended generation.
How it compares
Beam search is deterministic and breadth-first; temperature sampling is stochastic and single-path.
Comparisons that include Beam Search
Related topics
Questions that mention this term
Related terms
API LLM
An LLM accessed through a provider API: pay per token, get the frontier model, hand over ops.
FlashAttention
A memory-aware attention kernel that's 2-4x faster than vanilla, with identical math.
GGUF
Self-contained binary format for quantized LLMs; the standard for llama.cpp / Ollama / LM Studio.
Greedy Decoding
At each step, pick the single highest-probability token. Fast and deterministic, but often loops.
Grouped-Query Attention (GQA)
Compromise between MHA and MQA: query heads share KV heads in groups, cutting KV cache by 4-8x.
Knowledge Distillation
Train a small student model to match a big teacher's outputs: cheap, fast inference with most of the quality.