Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Attention Mechanism
/
Transformers
Transformers
Subtopic
47 questions
Questions tagged with Transformers — part of Attention Mechanism.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Where does cross-attention live in OpenAI Whisper?
Multiple Choice
Medium
Tying K and V to one projection, what is gained and what is risked?
Multiple Choice
Medium
Dividing pre-softmax attention scores by an extra factor > 1 at inference does what?
Multiple Choice
Medium
Define T5's relative position bias and what bucketing buys you
Short Answer
Medium
Walk through KV-cache updates when speculative decoding rejects some draft tokens
Short Answer
Medium
Along which axis of the QK^T score matrix is softmax applied inside attention?
Multiple Choice
Easy
Premium
Explain effective context for…
Short Answer
Medium
How does softmax turn an attention score into an attention weight?
Flashcard
Easy
Name the scalar that QK^T is divided by inside scaled dot product attention.
Flashcard
Easy
RMSNorm versus LayerNorm, what is kept and what is dropped?
Multiple Choice
Medium
Name what wraps the attention sub-layer in every transformer block.
Flashcard
Easy
Contrast absolute and relative positional encoding by what the score depends on
Multiple Choice
Medium
Fill the blank: for a single head, QK^T has shape (T_query, ___).
Fill in Blank
Easy
Premium
Why does normalizing Q…
Short Answer
Medium
Pre-norm versus post-norm: which placement makes deep stacks stable?
Multiple Choice
Medium
A batch packs a causal LM with right padded sequences and applies only the causal mask. Spot the mistake.
Spot the Error
Medium
What does a padding mask zero out, and why is it needed when batching variable length sequences?
Flashcard
Easy
Spot the masking bug in this packed sequence training setup
Spot the Error
Medium
Given Q of shape (B, n_heads, T, d_head), the per head attention output before concatenation has shape ___.
Fill in Blank
Easy
Describe the W_O projection in multi-head attention, its shape and what it mixes.
Flashcard
Easy
Beyond shape preservation, name two things W_O does after head concat
Short Answer
Medium
Which config token names the count of parallel attention heads in a layer?
Flashcard
Easy
Spot the flaw: a claim that attention by itself notices position.
Spot the Error
Medium
Premium
True or false: swapping…
Multiple Choice
Medium
Premium
Why does MQA underperform…
Short Answer
Medium
Showing 1–25 of 47
← Prev
Next →
Descript
Figure Ai
Cerebras
Sap
Databricks
Ola
Lepton Ai
Sierra
Cursor
Observe Ai
Mercor
Microsoft
Baidu
Ola
Capgemini
Cursor
Nykaa
Observe Ai
Ai4bharat
Cred
Reliance Jio
Sharechat
Coreweave
Doordash
Canva
Databricks
Copy Ai
Datarobot
Browserbase
Citadel
Ai21
Anduril
Cognizant
Deepseek
Elevenlabs
Freshworks
Cognizant
LangChain
Lepton Ai
Pinterest
Ai4bharat
Autodesk
Fiddler Ai
Mistral AI
Ai4bharat
Harvey
Bain
Niki Ai
Databricks
Jasper