Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM Evaluation
/
Faithfulness
Faithfulness
Subtopic
23 questions
Questions tagged with Faithfulness — part of LLM Evaluation.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
A RAG pipeline just shipped for internal docs. How would the team know if retrieval is actually helping?
Short Answer
Medium
A colleague claims their model is 'grounded' because it cites sources. Explain why citation presence alone does not prove factual grounding.
Multiple Choice
Medium
Describe a concrete, production runnable mechanism to detect hallucinations in a RAG answer (claims unsupported by the retrieved chunks), as the answer is generated or just after.
Short Answer
Medium
Match each RAGAS metric to what it specifically measures about a RAG pipeline.
Match Pairs
Medium
Premium
Your RAG system's faithfulness…
Multiple Choice
Medium
Complete the claim: the two axis decomposition of RAG eval that diagnoses which layer failed.
Fill in Blank
Medium
Which statements correctly describe the four RAGAS metrics?
Multi-select
Medium
Match each RAGAS metric to the specific RAG quality problem it detects
Match Pairs
Medium
What four metrics does the RAGAS paper propose and what does each measure?
Flashcard
Medium
Explain faithfulness vs answer relevance in RAG evaluation with concrete examples
Short Answer
Medium
Define faithfulness in RAG evaluation and distinguish it from answer relevance
Flashcard
Easy
What does context recall measure in RAG eval and what ground truth does it require?
Multiple Choice
Medium
What does context precision measure in RAG evaluation and what does low context precision indicate?
Multiple Choice
Medium
Predict what slice level eval reveals about an 85% faithful RAG system suspected to fail on multi-hop queries
Predict Output
Hard
Describe two automated hallucination detection techniques and their tradeoffs
Short Answer
Medium
Premium
What is the main…
Multiple Choice
Medium
Compare Promptfoo, DeepEval, and LangSmith as LLM eval frameworks: when to use each?
Short Answer
Hard
Clarify the relationship between unfaithfulness and hallucination in RAG systems
Short Answer
Hard
Premium
Is 'unfaithful' the same…
Multiple Choice
Medium
Explain FActScore's approach to long form factual evaluation and its design decisions
Short Answer
Hard
What does FActScore measure and how does it decompose long form factual evaluation?
Flashcard
Medium
Which of these are valid concerns when using LLM-as-judge for evaluation?
Multi-select
Medium
Which metric best measures whether a RAG answer is grounded in the retrieved context?
Multiple Choice
Medium
Gnani
Together Ai
Jump Trading
Perplexity
Goldman Sachs
Modal Labs
Anthropic
Coinbase
OpenAI
Pwc
Anduril
Mu Sigma
Anthropic
Banana Dev
Anthropic
Lakera
Bain
C3 Ai
Deloitte
Humanloop
Midjourney
Stripe
Anthropic
Cred
Baidu
Ey
Datadog
IBM
Browserbase
Runway
Lakera
Pinecone
Contextual Ai
Mphasis
Krutrim
Pinecone
Airbnb
Labelbox
Jump Trading
Lightning Ai
Hebbia
Mongodb
Sharechat
Snap
Descript
Groq