Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Agent Eval
Agent Eval
Subtopic
5 questions
Questions tagged with Agent Eval — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Design a diagnostic eval for an agentic workflow with 62% task success across 5 tools and 10+ steps that reveals WHY tasks fail, not just whether they pass.
Short Answer
Hard
Flashcard: what is an agent 'trajectory' and why does it matter for evaluation?
Flashcard
Easy
What structural property of SWE-bench makes it a harder eval than HumanEval for code agents?
Multiple Choice
Medium
Why is final task success rate insufficient for evaluating LLM agents, and what does trajectory eval add?
Short Answer
Hard
What does trajectory evaluation add beyond task success rate for LLM agent evaluation?
Multiple Choice
Medium
Haptik
Hebbia
Anthropic
Ey
Fractal Analytics
Niki Ai
Anthropic
OpenAI
Ai21
Ironclad