Agentic AI Developer
Claude · MCP · RAG · Spec-Kit
Initializing Agentic AI...
0%
Ramesha Javed
Founder & CEO, VisionDX AI
Building a RAG system that works in a demo is easy. Building one that holds up in production — with real users, diverse queries, and latency requirements — is a completely different challenge. This guide covers everything I've learned deploying RAG at scale.
The typical RAG demo uses a small, clean dataset with simple queries. Production is messier: users ask ambiguous questions, documents have inconsistent formatting, and you're dealing with thousands of chunks across multiple collections. The naive approach — chunk everything, embed it, retrieve top-k — breaks down fast.
Fixed-size chunking (512 tokens, overlap 50) is the default everyone starts with. It's wrong for most use cases. Semantic chunking — splitting on meaningful boundaries — consistently outperforms fixed-size by 15-30% in retrieval accuracy. For code documentation, I use AST-aware chunking. For PDFs with tables, I treat each table as an atomic chunk.
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
separators=["\n\n", "\n", ". ", " ", ""]
)
chunks = splitter.split_documents(docs)OpenAI text-embedding-3-small is the default choice, but it's not always optimal. For domain-specific content (medical, legal, code), fine-tuned models outperform general-purpose embeddings. I've had great results with nomic-embed-text for general use — it's 8x cheaper than OpenAI and within 2% accuracy on most benchmarks.
After testing Pinecone, Weaviate, and Qdrant, I settled on Qdrant for production. The filtering capabilities are unmatched — you can combine semantic search with metadata filters without a performance hit. Payload indexing lets you filter by date, category, or any metadata field at query time.
from qdrant_client import QdrantClient
from qdrant_client.models import Filter, FieldCondition, MatchValue
client = QdrantClient(url="http://localhost:6333")
results = client.search(
collection_name="knowledge_base",
query_vector=query_embedding,
query_filter=Filter(
must=[FieldCondition(key="category", match=MatchValue(value="technical"))]
),
limit=10
)Retrieval gets you candidates. Reranking gets you the right answer. Adding a cross-encoder reranker (Cohere Rerank or BGE-Reranker) after initial retrieval consistently improves answer quality by 20-40%. The cost is worth it for any production system. Retrieve top-50, rerank to top-5, then pass to LLM.
You can't improve what you don't measure. I use RAGAS (Retrieval Augmented Generation Assessment) for automated evaluation. Key metrics: faithfulness (is the answer grounded in retrieved context?), answer relevance (does it answer the question?), and context precision (are retrieved chunks relevant?). Run evaluation on a held-out test set before every deployment.