RoPE (Rotary Position Embedding)
Also known as: Rotary Position Embedding, Rotary embeddings
Position info injected by rotating Q and K vectors, easy to extend to longer contexts.
A positional encoding scheme that applies position-dependent rotations to query and key vectors in self-attention. Encodes relative position naturally and extends gracefully to longer sequences via scaling or interpolation.
In practice
Used in LLaMA, GPT-NeoX, Qwen, and most modern open LLMs. Senior interviews probe RoPE scaling (NTK, YaRN) for long-context extension.
Related topics
Questions that mention this term
Related terms
Attention Mechanism
How a model decides which input tokens to weight when computing each output token.
Causal Mask
Attention mask that hides future tokens so each position can only see itself and prior tokens.
Context Window
The max number of tokens a model can attend to at once.
Decoder-Only
Single autoregressive transformer stack: the shape of every modern frontier LLM.
Encoder-Decoder
Transformer with separate encoder + decoder stacks; strong for translation and structured seq2seq tasks.
FlashAttention
A memory-aware attention kernel that's 2-4x faster than vanilla, with identical math.