Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Roofline Model
Roofline Model
Subtopic
25 questions
Questions tagged with Roofline Model — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
TPOT is high, so the team plans to upgrade A100 to a faster compute GPU, critique
Spot the Error
Medium
Find the wrong move in 'decode is slow, so let's switch to a smaller FLOP model'
Spot the Error
Medium
Match Hopper FP8 and Blackwell FP4 tensor cores to their throughput and bandwidth effects on decode
Match Pairs
Medium
What unit measures GPU memory bandwidth and why does that number cap decode speed?
Flashcard
Easy
What does arithmetic intensity actually measure and what does AI near 1 imply?
Flashcard
Easy
Match each weight/activation quant regime to its mechanism and best fit workload
Match Pairs
Hard
Apply the roofline model to LLM inference and identify where decode and prefill sit on it.
Short Answer
Hard
On the roofline, which intervention moves an LLM decode workload most directly toward higher peak throughput?
Multiple Choice
Hard
Why does prefill saturate compute while decode is bottlenecked on memory bandwidth?
Short Answer
Medium
Premium
Which statement best explains…
Multiple Choice
Medium
Why does prefill's O(seq^2) attention cost saturate compute rather than bandwidth?
Short Answer
Hard
Why is the vision encoder phase a different bottleneck class from LLM decode in a multi-modal model?
Short Answer
Hard
Why is HBM bandwidth the binding constraint for LLM decode and how do you reason about the 'memory wall'?
Short Answer
Hard
Spot the errors in this 'optimize decode by reducing FLOPs' proposal
Spot the Error
Hard
Premium
Why is 'reduce FLOPs…
Short Answer
Medium
Predict the bandwidth vs compute latency of a single Llama-70B decode step on H100
Predict Output
Hard
Predict the critical batch size where decode crosses from memory bound to compute bound on H100
Predict Output
Hard
Why is batching the single biggest throughput lever for LLM decode?
Short Answer
Medium
Premium
What is the primary…
Multiple Choice
Medium
As batch size grows, where does throughput stop increasing and per request latency start exploding?
Short Answer
Hard
Predict the throughput and per request latency behaviour as batch size grows past saturation
Predict Output
Hard
Spot the errors in this 'batching is universal' claim
Spot the Error
Medium
Why does batching help LLM decode disproportionately more than batching helps CNN inference?
Short Answer
Medium
What is the arithmetic intensity of an LLM decode step and why is it close to 1?
Short Answer
Hard
Predict the arithmetic intensity of decode at varying batch sizes
Predict Output
Hard
Ironclad
Snap
Cred
Jasper
Modal Labs
NVIDIA
Mercor
NVIDIA
Flowise
NVIDIA
Flipkart
Graphcore
Fireworks Ai
Jpmorgan
Kpmg
Lepton Ai
Meesho
NVIDIA
Cognizant
Modal Labs
Flipkart
Lyzr
Shopify
Synthesia
NVIDIA
Turing
Ada
Cognizant
Netflix
Niki Ai
NVIDIA
Sigmoid
Ey
Gong
NVIDIA
Pinterest
Neptune Ai
NVIDIA
Bain
NVIDIA
Moveworks
NVIDIA
Airbnb
Graphcore
Decagon
Deloitte
Intuit
NVIDIA
Descript
Jump Trading