Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Attention Mechanism
/
Streamingllm
Streamingllm
Subtopic
8 questions
Questions tagged with Streamingllm — part of Attention Mechanism.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Pair KV-cache eviction and KV quantization with the situation each one is the right answer for
Match Pairs
Medium
Which token most often serves as an attention sink in pretrained autoregressive LLMs?
Multiple Choice
Easy
Why not just retrain models to eliminate attention sinks entirely?
Short Answer
Medium
Name two production KV-cache eviction policies and what each prioritizes
Short Answer
Medium
Find the bug: a Llama deployment serves traffic without prepending the BOS token.
Spot the Error
Medium
Name two recovery moves a serving stack can pull when KV memory runs short mid request.
Short Answer
Medium
Spot the errors in this KV eviction strategy for long running generation
Spot the Error
Hard
Why does naive sliding window KV eviction break and how does StreamingLLM's attention sink fix it?
Short Answer
Hard
Fiddler Ai
NVIDIA
Adobe
Freshworks
Flowise
Infosys
Notion
Swiggy
Cognizant
Dify
Microsoft
Pinterest
Oracle
Robinhood
Locus
Siemens