Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
AI Agents
/
Swe Bench
Swe Bench
Subtopic
8 questions
Questions tagged with Swe Bench — part of AI Agents.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Design a diagnostic eval for an agentic workflow with 62% task success across 5 tools and 10+ steps that reveals WHY tasks fail, not just whether they pass.
Short Answer
Hard
Evaluating a code generation model on real world tasks beyond HumanEval: which metrics cover correctness, efficiency, and style?
Multiple Choice
Medium
How should you score a multi-step agent where final answer accuracy is 15% but most steps are correct?
Multiple Choice
Medium
Walk through SWE-bench, AgentBench, GAIA, and TAU-bench: what they measure, their shared blind spots, and how to interpret agent benchmark numbers.
Short Answer
Hard
What structural property of SWE-bench makes it a harder eval than HumanEval for code agents?
Multiple Choice
Medium
What does SWE-bench measure that AgentBench does not, and why are both needed?
Short Answer
Hard
Match each agent benchmark to what it primarily measures
Match Pairs
Medium
How does a coding agent's scope differ from inline code completion?
Short Answer
Medium
Haptik
Hebbia
Contextual Ai
Haptik
Accenture
Cursor
C3 Ai
Cerebras
Cognizant
Dataiku
Capgemini
Databricks
Anthropic
OpenAI
Baseten
Cursor