Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
AI Agents
/
Agent Evaluation
Agent Evaluation
Subtopic
6 questions
Questions tagged with Agent Evaluation — part of AI Agents.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
How should you score a multi-step agent where final answer accuracy is 15% but most steps are correct?
Multiple Choice
Medium
What structural property of SWE-bench makes it a harder eval than HumanEval for code agents?
Multiple Choice
Medium
What does SWE-bench measure that AgentBench does not, and why are both needed?
Short Answer
Hard
Match each agent benchmark to what it primarily measures
Match Pairs
Medium
Premium
Why is scoring only…
Short Answer
Hard
What does trajectory evaluation measure that final answer accuracy alone cannot capture?
Multiple Choice
Medium
Capgemini
Databricks
Anthropic
OpenAI
Databricks
Datarobot
Accenture
Cursor
Databricks
Moveworks
Baseten
Cursor