Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Eval Frameworks
Eval Frameworks
Subtopic
5 questions
Questions tagged with Eval Frameworks — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Build an eval framework from scratch for three LLM products (support bot, code assistant, summarizer). Cover architecture, shared infra, per product customization, and six month failure modes.
Short Answer
Hard
HELM from Stanford claims to be a 'holistic' evaluation. What makes it different from running MMLU alone?
Multiple Choice
Easy
You hear about 'eval harnesses' like Promptfoo and DeepEval. Explain what an eval harness does that a Jupyter notebook cannot.
Flashcard
Easy
Compare Promptfoo, DeepEval, and LangSmith as LLM eval frameworks: when to use each?
Short Answer
Hard
Match each LLM eval framework to its primary differentiating capability
Match Pairs
Medium
Amd
OpenAI
OpenAI
Pwc
Elastic
Mphasis
Intel
Neptune Ai
Browserbase
Ltimindtree