Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
KV Cache
KV Cache
Subtopic
81 questions
Questions tagged with KV Cache — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
What PagedAttention borrows from operating system virtual memory
Flashcard
Easy
For 1M-token context decode, which parallelism axis actually relieves KV pressure?
Multiple Choice
Medium
Heavy multi-turn chat shares system prompts across requests, pick the serving framework that exploits that best.
Multiple Choice
Medium
At batch 16 and 4k context with FP8 weights, which model still fits on one 80 GB H100?
Multiple Choice
Medium
Per token decode time grows linearly as a chat session lengthens, debug it
Short Answer
Medium
A teammate randomized system prompts per user and expects the prompt cache to still help, what's wrong?
Spot the Error
Medium
Spot the error: 'we autoscale our LLM serving pods on CPU utilization.'
Spot the Error
Medium
Walk through how a prefix cache hit changes the work that chunked prefill has to do
Short Answer
Medium
Which quantization combo squeezes the most decode bandwidth per percent of quality lost?
Multiple Choice
Medium
Rank these chat API cost levers from highest to lowest typical ROI
Order Steps
Medium
Memory creeps up for hours then OOMs at 3am: diagnose the KV-cache leak pattern
Short Answer
Medium
Name the dominant cost driver of GPT-5.5 latency at a 200k-token input
Short Answer
Medium
Each decode step loads ___ from HBM, regardless of how many tokens have already been generated
Fill in Blank
Easy
Pair KV-cache eviction and KV quantization with the situation each one is the right answer for
Match Pairs
Medium
Sequence length vs context window, what is the practical distinction and which one drives KV cache size?
Flashcard
Easy
Prompt caching shows up on Anthropic and OpenAI pricing pages, what does it actually do?
Flashcard
Easy
Walk through what actually happens during the prefill phase of an LLM forward pass.
Flashcard
Easy
Why do hosted LLM APIs charge separate per million rates for input and output tokens?
Flashcard
Easy
Describe what happens during the decode phase of LLM inference.
Flashcard
Easy
What is the context window of an LLM?
Flashcard
Easy
Recognize vLLM and the optimization that made it famous
Flashcard
Easy
Identify PagedAttention and the project it ships with
Flashcard
Easy
Expand MQA and state its KV-cache cost tradeoff
Flashcard
Easy
Pair HBM and SRAM with their roles on an inference GPU
Match Pairs
Easy
Define GQA in transformer attention
Flashcard
Easy
Showing 1–25 of 81
← Prev
Next →
Cerebras
Gong
Cred
Flipkart
Palantir
Robinhood
Canva
Induced Ai
Browserbase
Ltimindtree
Anduril
Datarobot
Locus
Tech Mahindra
Hcl
Neo4j
Bytedance
Goldman Sachs
Amd
Meesho
Jasper
Niki Ai
Databricks
Dust
Fiddler Ai
Qdrant
NVIDIA
Qualcomm
Descript
NVIDIA
Cursor
Ey
Flowise
Krutrim
Airbnb
Capgemini
Accenture
Moveworks
Browserbase
Cognizant
Infosys
Tcs
Comet Ml
Persistent
Meesho
Robust Intelligence
Fiddler Ai
NVIDIA
Jasper
Servicenow