Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Batching
Batching
Subtopic
20 questions
Questions tagged with Batching — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Walk through how CUDA graphs reshape per step decode latency and when the win disappears
Short Answer
Medium
In benchmarks, tokens per second measures what exactly?
Flashcard
Easy
Throughput, for an LLM serving system: give the working definition
Flashcard
Easy
What does a padding mask zero out, and why is it needed when batching variable length sequences?
Flashcard
Easy
You pad a batch to equal length. How does the model know to ignore the filler tokens?
Flashcard
Easy
Premium
Spot the errors in…
Spot the Error
Medium
Apply the roofline model to LLM inference and identify where decode and prefill sit on it.
Short Answer
Hard
Premium
Why are output tokens…
Short Answer
Medium
Spot the errors in this 'optimize decode by reducing FLOPs' proposal
Spot the Error
Hard
Walk through the byte accounting that proves a single batch decode step is bandwidth bound on H100.
Short Answer
Hard
Premium
Why did continuous batching…
Short Answer
Medium
Predict the critical batch size where decode crosses from memory bound to compute bound on H100
Predict Output
Hard
Why is batching the single biggest throughput lever for LLM decode?
Short Answer
Medium
Premium
What is the primary…
Multiple Choice
Medium
As batch size grows, where does throughput stop increasing and per request latency start exploding?
Short Answer
Hard
Predict the throughput and per request latency behaviour as batch size grows past saturation
Predict Output
Hard
Spot the errors in this 'batching is universal' claim
Spot the Error
Medium
Why does batching help LLM decode disproportionately more than batching helps CNN inference?
Short Answer
Medium
What is the arithmetic intensity of an LLM decode step and why is it close to 1?
Short Answer
Hard
Predict the arithmetic intensity of decode at varying batch sizes
Predict Output
Hard
NVIDIA
Sierra
Meesho
NVIDIA
Anthropic
Cognizant
Intuit
NVIDIA
NVIDIA
Uber
Netflix
Niki Ai
NVIDIA
Sigmoid
Ey
Gong
NVIDIA
Pinterest
Arize Ai
Datadog
Braintrust
Deloitte
Elevenlabs
Freshworks
Rephrase Ai
Synthesia
Alibaba
Contextual Ai
Modal Labs
NVIDIA
Mercor
NVIDIA
Decagon
Deloitte
Intuit
NVIDIA
Descript
Jump Trading
Moveworks
NVIDIA