Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Time per Output Token
Time per Output Token
Subtopic
13 questions
Questions tagged with Time per Output Token — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Per token decode time grows linearly as a chat session lengthens, debug it
Short Answer
Medium
TPOT is high, so the team plans to upgrade A100 to a faster compute GPU, critique
Spot the Error
Medium
Find the wrong move in 'decode is slow, so let's switch to a smaller FLOP model'
Spot the Error
Medium
TTFT: name the metric it captures and what dominates it
Flashcard
Easy
In benchmarks, tokens per second measures what exactly?
Flashcard
Easy
End to end latency: fill in the queue, prefill and decode components.
Fill in Blank
Easy
Match TTFT and TPOT to the latency they measure during streaming
Match Pairs
Easy
Premium
Define TTFT, TPOT and…
Short Answer
Medium
A chatbot team reports TTFT P99 jumped from 300ms to 1.4s overnight. Which root cause is most likely?
Multiple Choice
Medium
What problem do DistServe and Splitwise solve by separating prefill and decode onto different GPUs?
Short Answer
Hard
Premium
For a 200-token prompt…
Short Answer
Medium
Premium
How does Sarathi-Serve's chunked…
Short Answer
Hard
Predict the throughput and per request latency behaviour as batch size grows past saturation
Predict Output
Hard
Cursor
Ey
Coinbase
NVIDIA
Canva
NVIDIA
Canva
Cred
Microsoft
NVIDIA
Ironclad
Snap
Cred
Jasper
Coreweave
Freshworks
Arize Ai
Datadog
Graphcore
Ltimindtree
Hcl
Meesho
Browserbase
NVIDIA
Intuit
NVIDIA