Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Attention Mechanism
/
Softmax
Softmax
Subtopic
28 questions
Questions tagged with Softmax — part of Attention Mechanism.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
Order the tail end…
Order Steps
Medium
Does asking for top-k logprobs add real compute cost per token, or only bytes on the wire?
Short Answer
Medium
How does the temperature parameter reshape the sampling distribution?
Flashcard
Easy
At which step in the attention pipeline does the mask actually take effect?
Multiple Choice
Easy
Each row of the post-softmax attention weight matrix corresponds to which side of the QK product?
Fill in Blank
Easy
Dividing pre-softmax attention scores by an extra factor > 1 at inference does what?
Multiple Choice
Medium
Pick what happens to attention weights when pre-softmax scores are divided by a temperature T > 1
Multiple Choice
Easy
Along which axis of the QK^T score matrix is softmax applied inside attention?
Multiple Choice
Easy
Name the two properties softmax guarantees for every row of the attention weight matrix
Fill in Blank
Easy
Which token most often serves as an attention sink in pretrained autoregressive LLMs?
Multiple Choice
Easy
Why not just retrain models to eliminate attention sinks entirely?
Short Answer
Medium
How does softmax turn an attention score into an attention weight?
Flashcard
Easy
Name the scalar that QK^T is divided by inside scaled dot product attention.
Flashcard
Easy
Predict what softmax produces when every key in a row is masked out.
Spot the Error
Medium
Where inside the attention sub-layer is dropout typically applied during training, and what does dropping there teach the model?
Short Answer
Medium
What does the dot product q . k actually measure inside an attention layer?
Flashcard
Easy
Identify the context vector inside scaled dot product attention and what produces it.
Multiple Choice
Easy
For a causal mask over a 4-token sequence, what value sits at row 2, column 3?
Predict Output
Easy
Attention temperature divides QK^T; sampling temperature divides output logits, distinguishwhere each lives.
Multiple Choice
Medium
Injecting noise into pre-softmax attention scores, goal and risk?
Multiple Choice
Medium
Your team debates 32K vs 128K vs 256K vocab. What is the core tradeoff they should frame?
Flashcard
Easy
Premium
What's the 'attention sink'…
Short Answer
Hard
Why is softmax used in attention rather than alternatives like sparsemax or simple sum normalization?
Multiple Choice
Hard
Complete the canonical attention formula: Attention(Q, K, V) = ___(QKᵀ / ___) · V
Fill in Blank
Easy
Where in attention is FP32 still used, and what breaks if you push everything to FP16/BF16?
Multiple Choice
Hard
Showing 1–25 of 28
← Prev
Next →
Bcg
Phonepe
Descript
OpenAI
Flowise
Infosys
Roblox
Robust Intelligence
Glean
Tata Digital
Decagon
Promptlayer
Browserbase
Citadel
Ai21
Anduril
Lepton Ai
Linkedin
Flipkart
Midjourney
Copy Ai
Sierra
Databricks
Ola
IBM
Mckinsey
Lepton Ai
Sierra
Adobe
Freshworks
Cursor
Swiggy
Bcg
Elevenlabs
Databricks
Induced Ai
Microsoft
Mu Sigma
Crewai
Fireworks Ai
Ey
Sharechat
N8n
Perplexity
Ey
Salesforce
Airbnb
Databricks
Droom
Pinterest