Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Text Gen Metrics
Text Gen Metrics
Subtopic
9 questions
Questions tagged with Text Gen Metrics — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
A colleague presents this…
Spot the Error
Medium
ROUGE keeps showing up in summarization evals. How does it work and where does it break?
Flashcard
Easy
BLEU score comes up in a meeting. What does it measure and when does it still make sense?
Flashcard
Easy
When is ROUGE-L more appropriate than ROUGE-2 for evaluating LLM outputs?
Multiple Choice
Medium
When should you use token level F1 vs exact match in a QA evaluation?
Multiple Choice
Medium
Explain the structural reason BLEU fails for open ended LLM evaluation
Short Answer
Medium
Why does BLEU score poorly predict human judgment for open ended LLM outputs?
Multiple Choice
Medium
Explain BERTScore's mechanism, its improvement over BLEU, and the specific failure mode that limits its use for factual LLM eval
Short Answer
Hard
How does BERTScore improve on BLEU, and what is its key limitation?
Multiple Choice
Medium
Fractal Analytics
Mistral AI
Airbnb
Fiddler Ai
Cognizant
Comet Ml
Elastic
Perplexity
Cursor
Kpmg
Elevenlabs
Salesforce
Neo4j
Pinterest
Cognizant
Hebbia
Browserbase
Intel