APEX Educational Institute

What is RAG? Build a Simple RAG Pipeline in Python (Step by Step)

Learn how Retrieval-Augmented Generation (RAG) works and build a small RAG pipeline in Python with chunking, embeddings, vector search and a grounded LLM answer.

Intermediate | 3 min read | Updated

Large Language Models (LLMs) are trained on public data up to a cut-off date. They do not know your company policies, your product manual or yesterday's support tickets. Retrieval-Augmented Generation (RAG) fixes this: before the model answers, you search your own documents and give the most relevant passages to the model as context.

How RAG works

  1. Ingest: load documents (PDF, web pages, wiki) and split them into small chunks.
  2. Embed: convert every chunk into a vector (a list of numbers that represents meaning).
  3. Store: save the vectors in a vector database or index.
  4. Retrieve: when a user asks a question, embed the question and find the closest chunks.
  5. Generate: send the question plus the retrieved chunks to the LLM and ask it to answer only from that context.
RAG does not retrain the model. You change the answer by changing the documents, which is cheaper, faster and easier to keep up to date than fine-tuning.

Step 1: Split documents into chunks

Chunks should be small enough to be specific but large enough to keep meaning. 300 to 800 characters with a small overlap is a good starting point.

python
def chunk_text(text: str, size: int = 500, overlap: int = 80) -> list[str]:
    chunks, start = [], 0
    while start < len(text):
        end = start + size
        chunks.append(text[start:end].strip())
        start = end - overlap
    return [c for c in chunks if c]

policy = open("leave_policy.txt", encoding="utf-8").read()
chunks = chunk_text(policy)
print(len(chunks), "chunks")

Step 2: Create embeddings

Any embedding model works. The example below uses the open-source sentence-transformers library, which runs locally for free.

python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
chunk_vectors = model.encode(chunks, normalize_embeddings=True)

normalize_embeddings=True makes every vector length 1, so a simple dot product equals cosine similarity.

Step 3: Retrieve the most relevant chunks

For a few thousand chunks, NumPy is enough. For millions, use a vector database such as pgvector, Qdrant or Pinecone.

python
import numpy as np

def retrieve(question: str, k: int = 3) -> list[str]:
    q = model.encode([question], normalize_embeddings=True)[0]
    scores = chunk_vectors @ q
    best = np.argsort(scores)[::-1][:k]
    return [chunks[i] for i in best]

context = retrieve("How many casual leaves do I get per year?")

Step 4: Generate a grounded answer

The prompt matters. Tell the model to use only the context and to say when it does not know.

python
def build_prompt(question: str, context: list[str]) -> str:
    joined = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(context))
    return f"""Answer the question using ONLY the context below.
If the answer is not in the context, say "I don't know based on the documents."
Cite the context numbers you used.

Context:
{joined}

Question: {question}"""

prompt = build_prompt("How many casual leaves do I get per year?", context)
# Send `prompt` to any LLM API (OpenAI, Gemini, Claude, or a local Llama model)

Common RAG mistakes

  • Chunks too large: the model gets noise and misses the key line.
  • No overlap: an answer split across two chunks is lost.
  • Only vector search: exact terms like product codes need keyword (BM25) search too. Combining both is called hybrid search.
  • No evaluation: keep a list of 20 to 50 real questions with expected answers and test every change against it.
  • Ignoring permissions: a user must only retrieve documents they are allowed to see.

Interview questions on RAG

QuestionShort answer
RAG vs fine-tuning?RAG adds knowledge at query time; fine-tuning changes model behaviour or style.
What is an embedding?A vector that represents meaning, so similar text gets nearby vectors.
How do you reduce hallucination?Better retrieval, strict prompts, citations and an "I don't know" fallback.
What is re-ranking?A second model that re-orders retrieved chunks by true relevance before generation.

Next steps

Try adding hybrid search, a re-ranker and source citations in the UI. To build production RAG apps with evaluation, guardrails and deployment, see our Generative AI Engineering course.

Master it hands-on

More Advanced AI tutorials