Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Fine-Tuning
/
Evaluation
Evaluation
Subtopic
29 questions
Questions tagged with Evaluation — part of Fine-Tuning.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Offline eval scores 0.9 but users report frequent wrong answers: why the gap, and what do you do?
Short Answer
Hard
Compute recall@5 and MRR for this labeled retrieval result
Predict Output
Medium
Select every metric that belongs on a serious multi-agent eval dashboard
Multi-select
Medium
How do you watch answer quality in production when nothing is labeled?
Short Answer
Hard
Why gate every prompt or model change behind a held out eval in CI?
Short Answer
Medium
What's the right way to roll out a prompt change you believe is better?
Multiple Choice
Medium
Walk through the triage order when a customer support LoRA tanks MMLU by 4 points post-deploy
Order Steps
Medium
Why does using the same model family as synthetic data teacher AND eval judge inflate FT scores?
Multiple Choice
Medium
Premium
When eval prompts appear…
Short Answer
Medium
A distilled model hits 92% on a public benchmark but 31% on an internal hold out: diagnose.
Short Answer
Medium
Picking a cheap but predictive eval to track during training: what do you actually log?
Short Answer
Medium
Premium
Train loss drops while…
Short Answer
Medium
A teammate asks what 'running evals' means in the LLM context. How would an engineer explain it without jargon?
Flashcard
Easy
How should you score a multi-step agent where final answer accuracy is 15% but most steps are correct?
Multiple Choice
Medium
Describe a concrete, production runnable mechanism to detect hallucinations in a RAG answer (claims unsupported by the retrieved chunks), as the answer is generated or just after.
Short Answer
Medium
Design a golden eval set for a production RAG system serving a legal research product. Explain how you build it, how big it should be, what each example contains, and what you measure with it.
Short Answer
Hard
Non-obvious modes of train/test leakage in instruction tuning evaluation
Short Answer
Hard
Which of these are real train/test leakage modes for an instruction tuning project?
Multi-select
Medium
Which of these belong on the ship checklist for a production fine-tune?
Multi-select
Medium
Design the post-FT eval suite for a customer support fine-tune
Short Answer
Hard
Which buckets MUST a post-FT eval suite include for a customer support fine-tune?
Multi-select
Medium
Premium
How do you detect…
Short Answer
Medium
Which signal is the most reliable indicator of catastrophic forgetting?
Multiple Choice
Medium
What is STRR and how does it complement fertility as a tokenizer fairness metric?
Flashcard
Medium
Match each RAGAS metric to what it specifically measures about a RAG pipeline.
Match Pairs
Medium
Showing 1–25 of 29
← Prev
Next →
Baidu
Doordash
Doordash
Meta
Ai4bharat
LangChain
Coinbase
LlamaIndex
Ai21
Bain
Browserbase
Haptik
Droom
OpenAI
Jump Trading
Perplexity
Dataiku
Replicate
Comet Ml
Polyai
Bcg
Groq
Ltimindtree
Midjourney
Cloudflare
Jump Trading
Contextual Ai
Qualcomm
Comet Ml
Tech Mahindra
Browserbase
Intuit
Persistent
Qualcomm
Ey
Sigmoid
Cerebras
Character Ai
Capgemini
Databricks
Banana Dev
Doordash
Hebbia
Ltimindtree
Humanloop
Stripe
LangChain
Turbopuffer
Browserbase
Runway