Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Benchmarks
Benchmarks
Subtopic
20 questions
Questions tagged with Benchmarks — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Measuring whether a vision-language model is actually good: what goes on the scorecard
Short Answer
Medium
Match each multimodal benchmark to the capability it actually measures
Match Pairs
Medium
A model aces a benchmark with suspiciously high scores. The first hypothesis is contamination. Explain what that means to a non-technical PM.
Flashcard
Easy
Every model release cites benchmark numbers. Define what a benchmark actually is and name two things it cannot tell you.
Flashcard
Easy
TruthfulQA is designed to catch a specific model failure. Describe what it tests and why larger models sometimes do worse.
Flashcard
Easy
MT-Bench is used alongside Chatbot Arena. Describe what MT-Bench tests that single turn benchmarks miss.
Flashcard
Easy
MMLU appears on every model leaderboard. Describe what it tests and its biggest blind spot.
Flashcard
Easy
HumanEval is the go to code generation benchmark. Describe the task it gives the model and how it decides if the answer is correct.
Flashcard
Easy
Chatbot Arena ranks models by Elo rating. Explain the Elo system to someone who has never played competitive chess.
Flashcard
Easy
Why does Chatbot Arena (LMSYS) give a different picture of model quality than static benchmarks like MMLU?
Multiple Choice
Easy
AlpacaEval 2 reports a 'win rate' for each model. Against what baseline and how is the winner decided?
Multiple Choice
Easy
Walk through SWE-bench, AgentBench, GAIA, and TAU-bench: what they measure, their shared blind spots, and how to interpret agent benchmark numbers.
Short Answer
Hard
What structural property of SWE-bench makes it a harder eval than HumanEval for code agents?
Multiple Choice
Medium
What makes multi-turn dialog evaluation harder than single turn, and how does MT-Bench address it?
Short Answer
Hard
Explain benchmark contamination and its effect on reported LLM capability scores
Short Answer
Medium
Which methods are practical approaches for detecting train eval data leakage in LLM benchmarks?
Multi-select
Hard
How does Chatbot Arena compute ELO rankings, and what bias affects its validity?
Short Answer
Hard
What is the primary validity threat to Chatbot Arena ELO rankings as a measure of general model quality?
Multiple Choice
Medium
Analyze MMLU's strengths and failure modes as an LLM benchmark
Short Answer
Hard
What are the strengths and weaknesses of MMLU as an LLM benchmark?
Multiple Choice
Medium
Accenture
Modal Labs
Contextual Ai
Haptik
Anthropic
Datarobot
LangChain
Shopify
Capgemini
OpenAI
Ltimindtree
OpenAI
Amd
Harvey
Amd
Snorkel Ai
Doordash
Uipath
Decagon
Inflection Ai
Capgemini
Krutrim
C3 Ai
Inflection Ai
Dust
Elastic
Ai4bharat
Autodesk
Dify
Notion
Airbnb
Bytedance
Anthropic
OpenAI
Elastic
Fireworks Ai
OpenAI
Weaviate
Citadel
LangChain