Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Home
/
Roles
/
AI Researcher
AI Researcher
621 questions
Works on the science, papers, theory, training dynamics, novel architectures, scaling laws.
Topics
Format
Difficulty
Companies
Track
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
Spot the bug in…
Spot the Error
Medium
Estimate the parameters weight tying saves a 128k-vocab model at d_model 2048
Multiple Choice
Easy
Walk through what quadrupling vocab to 128k does to params, compute, and sequences
Short Answer
Medium
Complete the path from token id to the vector that enters block 1
Fill in Blank
Easy
Derive the SwiGLU hidden width that matches a 4x GELU FFN parameter budget
Predict Output
Medium
Premium
Pinpoint the mechanism that…
Multiple Choice
Medium
Rescue a streaming deployment that collapses once early tokens leave the cache
Short Answer
Hard
Reason about KV cache growth in a sliding window model pushed to 32k tokens
Multiple Choice
Medium
Architect a layer pattern mixing sliding window and global attention for huge contexts
Short Answer
Hard
Explain why raising the RoPE base theta helps a model accept longer contexts
Multiple Choice
Medium
Pick the real reason Llama class models swapped LayerNorm for RMSNorm
Multiple Choice
Easy
Premium
Explain how a layer…
Multiple Choice
Medium
Spot the error in this transformer block that silently dropped its residuals
Spot the Error
Easy
Define QK-norm and the training failure it prevents
Flashcard
Medium
Diagnose attention logit explosion in a large run and defend QK-norm as the fix
Short Answer
Hard
Spot the error in this description of pre-norm placement and the final norm
Spot the Error
Medium
Diagnose why a 48-layer model diverges right after a pre-norm to post-norm refactor
Short Answer
Medium
Choose the positional scheme that fails hardest when context quadruples at inference
Multiple Choice
Medium
Identify what changes when a block computes attention and FFN in parallel, PaLM style
Multiple Choice
Medium
Premium
Debug a 64-layer stack…
Short Answer
Medium
Interpret what the logit lens reveals when you unembed intermediate layers
Multiple Choice
Hard
Premium
Select every architecture choice…
Multi-select
Medium
Calculate the per token KV cache footprint of a Llama-3-8B style config in fp16
Predict Output
Hard
Complete the per token KV cache memory formula
Fill in Blank
Medium
Identify which tensors your inference server actually stores in the KV cache
Multiple Choice
Easy
Showing 1–25 of 621
← Prev
Next →
Cerebras
Jpmorgan
Evenup
Goldman Sachs
Datarobot
Reliance Jio
Krutrim
Mphasis
Persistent
Polyai
Baseten
Jump Trading
Evenup
Infosys
Elevenlabs
Hcl
Ai4bharat
Neptune Ai
Induced Ai
Perplexity
Anyscale
Goldman Sachs
Banana Dev
Canva
Coreweave
Harvey
Descript
Jpmorgan
Anyscale
Coreweave
Samsung
Shield Ai
Fractal Analytics
Mercor
Bcg
Notion
Jane Street
Robust Intelligence
Palantir
Patronus
Ltimindtree
Modal Labs
Mercor
Mistral AI
Doordash
Mistral AI
Coreweave
Tencent
Fractal Analytics
Hebbia