Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Attention Mechanism
/
Grouped-Query Attention
Grouped-Query Attention
Subtopic
23 questions
Questions tagged with Grouped-Query Attention — part of Attention Mechanism.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Pick the recipe a 2026 frontier open weight LLM block actually ships
Multiple Choice
Medium
At batch 16 and 4k context with FP8 weights, which model still fits on one 80 GB H100?
Multiple Choice
Medium
Per token decode time grows linearly as a chat session lengthens, debug it
Short Answer
Medium
Expand MQA and state its KV-cache cost tradeoff
Flashcard
Easy
Define GQA in transformer attention
Flashcard
Easy
Tying K and V to one projection, what is gained and what is risked?
Multiple Choice
Medium
Spot the GQA configuration error: num_attention_heads=32, num_key_value_heads=7.
Spot the Error
Medium
Which config token names the count of parallel attention heads in a layer?
Flashcard
Easy
Premium
Why does MQA underperform…
Short Answer
Medium
Llama-2 70B uses 64 query heads, how many KV heads does it actually keep?
Multiple Choice
Easy
GQA, decode the acronym and describe its KV sharing pattern
Flashcard
Easy
Match GPT-2 versus Llama-2 attention design choices to their differences
Match Pairs
Medium
Decode phase attention is memory bandwidth bound. Explain what flips the regime.
Short Answer
Medium
How does DeepSeek's Multi-Latent Attention (MLA) compress the KV cache below GQA?
Short Answer
Hard
Match MLA, GQA and MQA to their mechanisms, cache reduction and quality posture
Match Pairs
Hard
Predict the KV cache size in GB for given model dimensions
Predict Output
Hard
Which statement most accurately captures why GQA, not MQA, became the production default?
Multiple Choice
Medium
Match each attention variant to its KV-cache property and a production user
Match Pairs
Medium
For a 64-Q-head model with GQA group size G=8, how many K/V heads exist and what is the cache reduction vs MHA?
Short Answer
Medium
Compare MHA, MQA, GQA, MLA, what production tradeoff are they all addressing, and which models use each?
Short Answer
Hard
Premium
Match each attention variant…
Match Pairs
Hard
Premium
Why is KV cache…
Short Answer
Hard
Compute the KV cache memory for a single request at 4096 context on Llama-2 70B (MHA) in FP16.
Predict Output
Hard
Figure Ai
Mercor
NVIDIA
Qualcomm
Cerebras
Sap
Droom
Mistral AI
Flowise
Meta
Cursor
Ey
Copy Ai
Datarobot
Harvey
Infosys
Accenture
Together Ai
Kpmg
Meta
Coreweave
Rephrase Ai
Cerebras
Kore Ai
Amd
Meesho
Jasper
Niki Ai
Lepton Ai
Pinterest
Accenture
Tesla
Bain
Cresta
Arize Ai
Hcl
Accenture
Bytedance
Canva
Graphcore
Alibaba
Anthropic
Coinbase
Modal Labs
Siemens
Together Ai