jnachi
Learning Hub
Agentic AI & RAG8 min readAdvanced

Automated RAG Evaluation: RAGAS & the TruLens RAG Triad

Implement continuous CI/CD evaluation for RAG pipelines using Context Relevance, Groundedness (Faithfulness), and Answer Relevance metrics.

Works with:RAGASTruLensDeepEvalOpenAI / Claude LLM-as-a-Judge

Key Takeaways

  • The RAG Triad evaluates the three critical edges: Context Relevance (Retriever), Groundedness (Generator), and Answer Relevance (User intent)
  • Groundedness / Faithfulness measures whether every assertion in the response is mathematically backed by the retrieved context
  • Context Relevance assesses the signal-to-noise ratio in retrieved passages to avoid paying for bloated context windows
  • RAGAS automated scoring pipelines run against golden evaluation datasets in GitHub Actions before deploying prompt or retriever changes

The Diagnostic Context

You cannot optimize what you do not measure. Guessing whether a RAG prompt change improved response accuracy leads to silent production degradation. Automated evaluation frameworks turn subjective reviews into repeatable unit test scores.

The Core Technique

The RAG Triad Evaluation Metrics

CODE / PROMPT
       [ User Query ]
          /        \
  Context          Answer
 Relevance        Relevance
      /              \
 [ Retrieved Context ] ---> [ Generated Answer ]
              \               /
               -- Faithfulness --
  1. Context Relevance: Does the retrieved chunk contain only information relevant to answering the query?
  2. Faithfulness / Groundedness: Are all statements in the generated response derived exclusively from the context (0% hallucination)?
  3. Answer Relevance: Does the response directly address the user's specific prompt without topic drift?
PYTHON
# Example RAGAS Evaluation Script
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision
from datasets import Dataset

eval_data = {
    "question": ["What is the maximum file upload limit on the Pro plan?"],
    "contexts": [["The Pro plan supports single file uploads up to 500 MB. Enterprise supports 5 GB."]],
    "answer": ["The Pro plan allows single file uploads up to 500 MB."],
    "ground_truth": ["500 MB per file on the Pro plan."]
}

dataset = Dataset.from_dict(eval_data)
results = evaluate(dataset, metrics=[faithfulness, answer_relevancy, context_precision])
print("Faithfulness Score:", results["faithfulness"]) # 1.0 = 100% grounded
5-Minute Activation Challenge

Try This Right Now

Write an evaluation test case with an intentional hallucination (e.g., context says 500 MB, answer says 2 GB) and verify that the faithfulness score drops to 0.0!

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (1 Questions)

1

If a RAG system generates an answer that is polite and fluent, but claims a policy detail not present in the retrieved context, which RAG Triad metric will score low?