Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Statistics
Statistics
Subtopic
11 questions
Questions tagged with Statistics — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
Your eval shows Model…
Multiple Choice
Medium
Explain the sign reversal failure mode in LLM judge calibration and how to prevent it
Short Answer
Hard
Premium
What is the sign…
Multiple Choice
Hard
How do you determine if a 2% eval score drop between model versions is statistically significant?
Short Answer
Hard
Which test is most appropriate for determining if a binary metric LLM eval drop is statistically significant?
Multiple Choice
Medium
Explain the pass@k metric and why the unbiased estimator matters for code gen evaluation
Short Answer
Hard
What does Cohen's kappa measure in human LLM evaluation?
Flashcard
Easy
How large should a golden eval set be, and what signals tell you when you have enough examples?
Short Answer
Hard
When should you use Fleiss's kappa instead of Cohen's kappa in a human eval study?
Multiple Choice
Medium
How do you measure whether an LLM judge is well calibrated against human raters?
Short Answer
Medium
Complete: Cohen's kappa >___ is 'substantial' agreement; >___ is 'almost perfect' agreement
Fill in Blank
Medium
Coinbase
Qualcomm
Banana Dev
Perplexity
Intuit
OpenAI
Gnani
Sarvam
Goldman Sachs
Redis
Graphcore
OpenAI
OpenAI
Sharechat
Cursor
Sarvam
Deepseek
Sourcegraph
Glean
Stripe
Stability Ai
Vellum