Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Fine-Tuning
/
Loss Masking
Loss Masking
Subtopic
15 questions
Questions tagged with Loss Masking — part of Fine-Tuning.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
Why does completion only…
Multiple Choice
Medium
Premium
Spot the bug: chat…
Spot the Error
Medium
Walk through one training example for SFT'ing a single weather tool call.
Short Answer
Medium
Should the system message contribute to the SFT loss?
Multiple Choice
Easy
Define supervised fine-tuning (SFT) in one breath.
Flashcard
Easy
`max_seq_len` controls one bound, what happens to examples that exceed it?
Flashcard
Easy
Which loss function does supervised fine-tuning actually minimize?
Multiple Choice
Easy
Why does dropping the EOS token from SFT labels produce a model that never stops generating?
Flashcard
Easy
DeepSeek style reasoning distillation: what travels from teacher to student, and what gets cut from the loss?
Short Answer
Medium
DataCollator pads to max_length and sets attention_mask = ones: find what breaks
Spot the Error
Medium
A batch packs a causal LM with right padded sequences and applies only the causal mask. Spot the mistake.
Spot the Error
Medium
Spot the error in this description of SFT loss masking
Spot the Error
Hard
Explain SFT's loss and the role of prompt token masking
Short Answer
Medium
Spot the bug in this sequence packing setup
Spot the Error
Hard
Spot the bug: a Llama-3 fine-tune that skips the end of turn token
Spot the Error
Hard
Canva
Lightning Ai
Phonepe
Pwc
Infosys
Moveworks
Alibaba
OpenAI
Datarobot
OpenAI
Ai4bharat
Harvey
Airbnb
Mercor
N8n
NVIDIA
OpenAI
Roblox
Voyage Ai
Yellow Ai
Glean
Meesho
Anduril
Lepton Ai
Gong
Stripe
Ai4bharat
Razorpay
Braintrust
Jane Street