Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Home
/
Search
Find questions
Search
Find questions across titles, topics, and companies.
Search
30 results for "transformers"
Premium
Which part of a…
Multiple Choice
Medium
Walk through how a token's d_model embedding is split into per head pieces inside multi-head attention.
Multiple Choice
Easy
Identify the context vector inside scaled dot product attention and what produces it.
Multiple Choice
Easy
Contrast ALiBi's linear distance bias with RoPE's vector rotation, which one modifies scores additively, and which one modifies Q and K directly?
Multiple Choice
Medium
Dividing pre-softmax attention scores by an extra factor > 1 at inference does what?
Multiple Choice
Medium
Identify the attention pattern difference between Mistral 7B and Llama-2 7B
Multiple Choice
Medium
Why is causal attention cheaper than full bidirectional at the same length?
Multiple Choice
Medium
Pre-norm versus post-norm: which placement makes deep stacks stable?
Multiple Choice
Medium
Which sublayer of a dense transformer block carries more parameters, attention or FFN?
Multiple Choice
Easy
Interpret what the logit lens reveals when you unembed intermediate layers
Multiple Choice
Hard
Premium
Explain how a layer…
Multiple Choice
Medium
Walk through a T5 decoder block in order, list the attention sub-layers and what each one's Q, K, V read from.
Order Steps
Medium
Order the pipeline that stretches an 8k pretrained model to 128k context
Order Steps
Hard
Premium
Order the operations inside…
Order Steps
Medium
Spot the masking bug in this packed sequence training setup
Spot the Error
Medium
Critique this claim about CLIP using cross-attention between towers
Spot the Error
Medium
Premium
Why did decoder-only architectures…
Short Answer
Medium
Premium
Explain effective context for…
Short Answer
Medium
Premium
Why does MQA underperform…
Short Answer
Medium
How does a Vision Transformer turn an image into a sequence, and what mask runs over it?
Short Answer
Medium
Premium
Why does normalizing Q…
Short Answer
Medium
Define T5's relative position bias and what bucketing buys you
Short Answer
Medium
Premium
When do you need…
Short Answer
Medium
Chinchilla's lesson: how did it reshape architecture and training choices after 2022?
Short Answer
Hard
Describe the W_O projection in multi-head attention, its shape and what it mixes.
Flashcard
Easy
Why Hugging Face Hub functions as the de facto model registry for open weights
Flashcard
Easy
Name the scalar that QK^T is divided by inside scaled dot product attention.
Flashcard
Easy
Fill the blank: for a single head, QK^T has shape (T_query, ___).
Fill in Blank
Easy
Match GPT-2 versus Llama-2 attention design choices to their differences
Match Pairs
Medium
Premium
Match each positional encoding…
Match Pairs
Medium
Graphcore
Mckinsey
Neo4j
Ola
Cursor
Swiggy
Datarobot
Elevenlabs
Databricks
Ola
Descript
Pinterest
Netflix
Typeface
Baidu
Ola
Databricks
Phonepe
Banana Dev
Canva
Ai4bharat
Neptune Ai
Cursor
Goldman Sachs
Haptik
Perplexity
Mistral AI
NVIDIA
Bain
Niki Ai
Polyai
Qdrant
Dataiku
OpenAI
Reliance Jio
Sharechat
Copy Ai
Datarobot
Coinbase
Meesho
Coreweave
Doordash
Nykaa
Observe Ai
Anduril
Robinhood
Anthropic
OpenAI
Cognizant
LangChain
Autodesk
Lepton Ai
Ai21
Anduril
Ai4bharat
Autodesk
Arize Ai
Hcl
Braintrust
Jane Street