Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Decode
Decode
Subtopic
25 questions
Questions tagged with Decode — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Find the wrong move in 'decode is slow, so let's switch to a smaller FLOP model'
Spot the Error
Medium
Which quantization combo squeezes the most decode bandwidth per percent of quality lost?
Multiple Choice
Medium
Name the dominant cost driver of GPT-5.5 latency at a 200k-token input
Short Answer
Medium
Each decode step loads ___ from HBM, regardless of how many tokens have already been generated
Fill in Blank
Easy
Prefix cache hit lands on a request: does speculative decoding still help that request?
Short Answer
Medium
Walk through how CUDA graphs reshape per step decode latency and when the win disappears
Short Answer
Medium
TTFT: name the metric it captures and what dominates it
Flashcard
Easy
The roofline model in plain language, without the chart, what is it saying?
Flashcard
Easy
What unit measures GPU memory bandwidth and why does that number cap decode speed?
Flashcard
Easy
End to end latency: fill in the queue, prefill and decode components.
Fill in Blank
Easy
Why do hosted LLM APIs charge separate per million rates for input and output tokens?
Flashcard
Easy
Describe what happens during the decode phase of LLM inference.
Flashcard
Easy
Match TTFT and TPOT to the latency they measure during streaming
Match Pairs
Easy
During decode, why is only Q computed for the new token while full K and V come from cache?
Multiple Choice
Easy
Prefill and decode: name the two phases of LLM inference and say which one is compute bound
Flashcard
Easy
Decode phase attention is memory bandwidth bound. Explain what flips the regime.
Short Answer
Medium
What problem does chunked prefill solve in production LLM serving?
Multiple Choice
Medium
Why does prefill saturate compute while decode is bottlenecked on memory bandwidth?
Short Answer
Medium
Premium
Which statement best explains…
Multiple Choice
Medium
What exactly does the KV cache store and what computational redundancy does it eliminate?
Short Answer
Medium
Premium
Why are output tokens…
Short Answer
Medium
Premium
For a 200-token prompt…
Short Answer
Medium
Walk through the byte accounting that proves a single batch decode step is bandwidth bound on H100.
Short Answer
Hard
Predict the bandwidth vs compute latency of a single Llama-70B decode step on H100
Predict Output
Hard
Spot the errors in this 'API is slow because of network and tokenizer' explanation
Spot the Error
Medium
Cred
Jasper
Ey
NVIDIA
Descript
NVIDIA
Autodesk
Modal Labs
Alibaba
Mistral AI
Bain
NVIDIA
Accenture
Moveworks
Ai21
Airbnb
NVIDIA
Sierra
Harvey
Infosys
Cognizant
Modal Labs
Inflection Ai
Lepton Ai
Anthropic
Cognizant
Canva
Cred
Intuit
NVIDIA
Meesho
Robust Intelligence
Graphcore
Ltimindtree
Coreweave
Freshworks
Jane Street
Qdrant
Fireworks Ai
Jpmorgan
Browserbase
Ltimindtree
Anduril
Datarobot
Character Ai
Deepseek
Hcl
Meesho
Airbnb
Graphcore