Agentic AI Developer
Claude · MCP · RAG · Spec-Kit
Initializing Agentic AI...
0%
Ramesha Javed
Founder & CEO, VisionDX AI
Most teams evaluate RAG by asking 'does it get the right answer?' That's necessary but nowhere near sufficient. Here's the complete evaluation framework I use in production.
RAGAS (Retrieval Augmented Generation Assessment) provides four core metrics: Faithfulness (is the answer supported by retrieved context?), Answer Relevance (does the answer address the question?), Context Precision (are the retrieved chunks actually relevant?), and Context Recall (were all relevant chunks retrieved?). You need all four — optimizing for one often hurts the others.
RAGAS uses Claude or GPT-4 as a judge, which means evaluation itself costs money. Be smart about when to run full evaluation — I run it nightly on a 100-sample benchmark, and on every deployment that changes the retrieval pipeline.
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision
dataset = Dataset.from_dict({
"question": questions,
"answer": answers,
"contexts": retrieved_contexts,
"ground_truth": ground_truths
})
result = evaluate(
dataset,
metrics=[faithfulness, answer_relevancy, context_precision]
)
print(result.to_pandas())Automated metrics catch regressions but miss nuance. I run monthly human evaluation sessions with domain experts — for VisionDX, that means radiologists reviewing AI explanations. Human evaluators catch hallucinations that automated metrics miss and identify systematic biases the model has developed.
The evaluation pipeline should be as automated as your CI/CD pipeline. Every PR that touches the RAG stack triggers evaluation. If faithfulness drops below 0.85, the build fails. This sounds strict but it's saved us from shipping several bad retrieval changes. Treat evaluation like tests — it's part of the definition of done.