Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Prefill
Prefill
Subtopic
22 questions
Questions tagged with Prefill — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Name the dominant cost driver of GPT-5.5 latency at a 200k-token input
Short Answer
Medium
Prefix cache hit lands on a request: does speculative decoding still help that request?
Short Answer
Medium
TTFT: name the metric it captures and what dominates it
Flashcard
Easy
The roofline model in plain language, without the chart, what is it saying?
Flashcard
Easy
Prompt caching shows up on Anthropic and OpenAI pricing pages, what does it actually do?
Flashcard
Easy
Walk through what actually happens during the prefill phase of an LLM forward pass.
Flashcard
Easy
End to end latency: fill in the queue, prefill and decode components.
Fill in Blank
Easy
Why do hosted LLM APIs charge separate per million rates for input and output tokens?
Flashcard
Easy
FLOPs vs FLOP/s: fill in the counts versus rate distinction.
Fill in Blank
Easy
Describe what happens during the decode phase of LLM inference.
Flashcard
Easy
Does Anthropic's message_start event count toward TTFT in production benchmarks?
Short Answer
Medium
Match TTFT and TPOT to the latency they measure during streaming
Match Pairs
Easy
Prefill and decode: name the two phases of LLM inference and say which one is compute bound
Flashcard
Easy
What problem does chunked prefill solve in production LLM serving?
Multiple Choice
Medium
How does prefix KV-cache sharing reduce serving cost across requests?
Short Answer
Medium
Which prompt construction pattern is most likely to actually hit the provider's prompt cache?
Multiple Choice
Medium
Why does prefill saturate compute while decode is bottlenecked on memory bandwidth?
Short Answer
Medium
Premium
Which statement best explains…
Multiple Choice
Medium
Why does prefill's O(seq^2) attention cost saturate compute rather than bandwidth?
Short Answer
Hard
Premium
Why are output tokens…
Short Answer
Medium
Why does FlashAttention speed up serving even though it doesn't change the math?
Short Answer
Hard
Where does FlashAttention deliver the biggest serving speedup?
Multiple Choice
Medium
Accenture
Moveworks
Ai21
Airbnb
Anthropic
Snap
Descript
Figure Ai
Cognizant
Modal Labs
Flipkart
Lyzr
Anthropic
Cognizant
Fireworks Ai
NVIDIA
Coreweave
Freshworks
Jane Street
Qdrant
Palantir
Robinhood
Canva
Induced Ai
Browserbase
Ltimindtree
Anduril
Datarobot
Character Ai
Deepseek
Graphcore
Ltimindtree
Ada
Comet Ml
Hcl
Meesho
Alibaba
Mistral AI
Anthropic
OpenAI
Bain
NVIDIA
Fireworks Ai
NVIDIA