Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Arithmetic Intensity
Arithmetic Intensity
Subtopic
25 questions
Questions tagged with Arithmetic Intensity — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
The roofline model in plain language, without the chart, what is it saying?
Flashcard
Easy
Walk through what actually happens during the prefill phase of an LLM forward pass.
Flashcard
Easy
What unit measures GPU memory bandwidth and why does that number cap decode speed?
Flashcard
Easy
Why do hosted LLM APIs charge separate per million rates for input and output tokens?
Flashcard
Easy
FLOPs vs FLOP/s: fill in the counts versus rate distinction.
Fill in Blank
Easy
What does arithmetic intensity actually measure and what does AI near 1 imply?
Flashcard
Easy
Apply the roofline model to LLM inference and identify where decode and prefill sit on it.
Short Answer
Hard
On the roofline, which intervention moves an LLM decode workload most directly toward higher peak throughput?
Multiple Choice
Hard
Why does prefill saturate compute while decode is bottlenecked on memory bandwidth?
Short Answer
Medium
Premium
Which statement best explains…
Multiple Choice
Medium
Why does prefill's O(seq^2) attention cost saturate compute rather than bandwidth?
Short Answer
Hard
Why is HBM bandwidth the binding constraint for LLM decode and how do you reason about the 'memory wall'?
Short Answer
Hard
Premium
Why are output tokens…
Short Answer
Medium
Spot the errors in this 'optimize decode by reducing FLOPs' proposal
Spot the Error
Hard
Premium
Why is 'reduce FLOPs…
Short Answer
Medium
Walk through the byte accounting that proves a single batch decode step is bandwidth bound on H100.
Short Answer
Hard
Predict the bandwidth vs compute latency of a single Llama-70B decode step on H100
Predict Output
Hard
Predict the critical batch size where decode crosses from memory bound to compute bound on H100
Predict Output
Hard
Why is batching the single biggest throughput lever for LLM decode?
Short Answer
Medium
Premium
What is the primary…
Multiple Choice
Medium
As batch size grows, where does throughput stop increasing and per request latency start exploding?
Short Answer
Hard
Spot the errors in this 'batching is universal' claim
Spot the Error
Medium
Why does batching help LLM decode disproportionately more than batching helps CNN inference?
Short Answer
Medium
What is the arithmetic intensity of an LLM decode step and why is it close to 1?
Short Answer
Hard
Predict the arithmetic intensity of decode at varying batch sizes
Predict Output
Hard
Jane Street
Qdrant
Canva
Induced Ai
Fireworks Ai
Jpmorgan
Browserbase
Ltimindtree
Kpmg
Lepton Ai
Ada
Comet Ml
Meesho
NVIDIA
Cognizant
Modal Labs
Flipkart
Lyzr
NVIDIA
Turing
Anthropic
Cognizant
Ada
Cognizant
Intuit
NVIDIA
Netflix
Niki Ai
NVIDIA
Sigmoid
Ey
Gong
NVIDIA
Pinterest
Neptune Ai
NVIDIA
Bain
NVIDIA
Moveworks
NVIDIA
Modal Labs
NVIDIA
Mercor
NVIDIA
Airbnb
Graphcore
Decagon
Deloitte
Descript
Jump Trading