Large Language Models (LLMs) are trained on public data up to a cut-off date. They do not know your company policies, your product manual or yesterday's support tickets. Retrieval-Augmented Generation (RAG) fixes this: before the model answers, you search your own documents and give the most relevant passages to the model as context.
How RAG works
- Ingest: load documents (PDF, web pages, wiki) and split them into small chunks.
- Embed: convert every chunk into a vector (a list of numbers that represents meaning).
- Store: save the vectors in a vector database or index.
- Retrieve: when a user asks a question, embed the question and find the closest chunks.
- Generate: send the question plus the retrieved chunks to the LLM and ask it to answer only from that context.
RAG does not retrain the model. You change the answer by changing the documents, which is cheaper, faster and easier to keep up to date than fine-tuning.
Step 1: Split documents into chunks
Chunks should be small enough to be specific but large enough to keep meaning. 300 to 800 characters with a small overlap is a good starting point.
def chunk_text(text: str, size: int = 500, overlap: int = 80) -> list[str]:
chunks, start = [], 0
while start < len(text):
end = start + size
chunks.append(text[start:end].strip())
start = end - overlap
return [c for c in chunks if c]
policy = open("leave_policy.txt", encoding="utf-8").read()
chunks = chunk_text(policy)
print(len(chunks), "chunks")Step 2: Create embeddings
Any embedding model works. The example below uses the open-source sentence-transformers library, which runs locally for free.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
chunk_vectors = model.encode(chunks, normalize_embeddings=True)normalize_embeddings=True makes every vector length 1, so a simple dot product equals cosine similarity.
Step 3: Retrieve the most relevant chunks
For a few thousand chunks, NumPy is enough. For millions, use a vector database such as pgvector, Qdrant or Pinecone.
import numpy as np
def retrieve(question: str, k: int = 3) -> list[str]:
q = model.encode([question], normalize_embeddings=True)[0]
scores = chunk_vectors @ q
best = np.argsort(scores)[::-1][:k]
return [chunks[i] for i in best]
context = retrieve("How many casual leaves do I get per year?")Step 4: Generate a grounded answer
The prompt matters. Tell the model to use only the context and to say when it does not know.
def build_prompt(question: str, context: list[str]) -> str:
joined = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(context))
return f"""Answer the question using ONLY the context below.
If the answer is not in the context, say "I don't know based on the documents."
Cite the context numbers you used.
Context:
{joined}
Question: {question}"""
prompt = build_prompt("How many casual leaves do I get per year?", context)
# Send `prompt` to any LLM API (OpenAI, Gemini, Claude, or a local Llama model)Common RAG mistakes
- Chunks too large: the model gets noise and misses the key line.
- No overlap: an answer split across two chunks is lost.
- Only vector search: exact terms like product codes need keyword (BM25) search too. Combining both is called hybrid search.
- No evaluation: keep a list of 20 to 50 real questions with expected answers and test every change against it.
- Ignoring permissions: a user must only retrieve documents they are allowed to see.
Interview questions on RAG
| Question | Short answer |
|---|---|
| RAG vs fine-tuning? | RAG adds knowledge at query time; fine-tuning changes model behaviour or style. |
| What is an embedding? | A vector that represents meaning, so similar text gets nearby vectors. |
| How do you reduce hallucination? | Better retrieval, strict prompts, citations and an "I don't know" fallback. |
| What is re-ranking? | A second model that re-orders retrieved chunks by true relevance before generation. |
Next steps
Try adding hybrid search, a re-ranker and source citations in the UI. To build production RAG apps with evaluation, guardrails and deployment, see our Generative AI Engineering course.
