Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Grouped Query Attention
Grouped Query Attention
Subtopic
5 questions
Questions tagged with Grouped Query Attention — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Spot the GQA configuration error: num_attention_heads=32, num_key_value_heads=7.
Spot the Error
Medium
Match MLA, GQA and MQA to their mechanisms, cache reduction and quality posture
Match Pairs
Hard
Compute the KV cache size for a 70B model at 128k context (80 layers, 64 heads, head_dim 128, FP16).
Short Answer
Hard
Match each attention variant to its KV-cache property and a production user
Match Pairs
Medium
For a 64-Q-head model with GQA group size G=8, how many K/V heads exist and what is the cache reduction vs MHA?
Short Answer
Medium
Bain
Cresta
Accenture
Bytedance
Canva
Graphcore
Intuit
NVIDIA
Kpmg
Meta