RAG Architecture & Workflow

Master the complete end-to-end blueprint of a RAG system—understanding the crucial separation between Offline Indexing (Load, Chunk, Embed, Store) and Online Querying & Generation.

28 minBeginnerCode Examples

The Core Idea: A beginner often thinks a RAG system reads and processes 1,000 PDFs from scratch every single time a user asks a question. If it did that, every answer would take 20 minutes! Instead, every production RAG system is split into two separate workflows: (1) the Offline Indexing Phase (which prepares, chops, embeds, and organizes your documents into a Vector Database ahead of time—like stocking a library's shelves), and (2) the Online Query Phase (which searches that pre-built index in milliseconds to answer a live user question).

  1. Beginner: The Two-Phase Secret (Stocking the Library vs. Answering a Visitor)

Imagine walking into a city library and asking the librarian: "Which book explains how solar panels work?"

The librarian does not open all 50,000 books on the floor and read them while you wait! Months before you arrived, the library ran an Indexing Process—cataloging every chapter by topic so the librarian can look up the exact shelf in 5 seconds.

A RAG System works the exact same way by separating its work into two phases:

Phase 1: Offline Ingestion & IndexingRuns Ahead of Time
When it runs: Once when you set up your app, or in the background whenever a new PDF/webpage is uploaded.
What it does: Loads raw files ➔ Splits them into chunks ➔ Converts chunks into Embedding vectors ➔ Saves them inside a Vector Database.
Phase 2: Online Retrieval & GenerationRuns Live per User
When it runs: In real time (under 1 second) every time a user types a question into the chat box.
What it does: Embeds the user's question ➔ Searches the Vector Database for the top matching chunks ➔ Sends Question + Chunks to the LLM ➔ Returns the final answer!

  1. Beginner: Step-by-Step Through Phase 1 (Offline Document Indexing)

How do we turn a messy folder of company PDFs, Word docs, and web pages into a searchable AI memory bank? We pass the documents through five sequential stations during Offline Indexing:

STAGE 1
Source Documents

Raw files: PDFs, DOCX, Markdown, HTML web pages, Notion wikis, or CSVs.

STAGE 2
Document Loading

Extracts clean text strings and attaches metadata (filename, page number, author).

STAGE 3
Chunking / Splitting

Slices long 50-page documents into bite-sized passages (e.g., 300 to 500 words each).

STAGE 4
Embedding Model

Converts each text chunk into a dense meaning vector (e.g., 1,536 numbers).

STAGE 5
Vector Database

Stores the vectors + original text + metadata together in a fast search index!

⚡ Knowledge Check

Why do we split a 100-page PDF into smaller chunks (Stage 3) before creating embeddings, instead of turning the entire 100-page book into one single vector?

A) One vector for 100 pages blurs all topics together, and we only want to pass the specific relevant paragraph to the LLM▼
✓ Correct!Embedding models have a token limit and work best on focused passages. Chunking lets us pinpoint the exact paragraph that answers the user's question!
B) Because hard drives cannot store files larger than 1 page▼
✕ Incorrect.Storage is not the issue—chunking is done to preserve fine-grained semantic search precision and keep retrieved prompt context concise.

  1. Medium: Step-by-Step Through Phase 2 (Online Retrieval & Generation)

Once the Vector Database is stocked and ready, what happens in the 500 milliseconds after a user types a question into your app? The Online Query Phase executes 5 live steps:

LIVE STEP 1 · USER QUERY
User submits a natural-language question

"How do I reset my corporate VPN password?"

LIVE STEP 2 · QUERY EMBEDDING
Embed the question using the EXACT SAME Embedding Model

Query Vector q = [0.14, -0.62, 0.88, ...]

LIVE STEP 3 · SIMILARITY SEARCH
Compare Query Vector against all Chunk Vectors in Vector DB

Cosine Similarity ➔ Top-K (e.g., Top 3 Chunks)

LIVE STEP 4 · PROMPT AUGMENTATION
Format System Instructions + Retrieved Chunks + User Question

Grounded Prompt Template

LIVE STEP 5 · LLM GENERATION
LLM reads the augmented prompt and streams the cited answer!

"Go to vpn.company.com and click..."

⚠️ The #1 Golden Rule of RAG: Never Mix Embedding Models!

Notice Live Step 2 above: the model that embeds the User Query in Phase 2 MUST be the exact same Embedding Model that embedded the Document Chunks in Phase 1! Different embedding models speak completely different mathematical coordinate languages—comparing a query vector from Model A to document vectors from Model B produces random garbage!

  1. Medium: Offline Indexing vs. Online Querying — Head-to-Head Comparison

Understanding the separation between Offline Indexing and Online Querying is essential for designing and debugging any RAG system:

System AspectPhase 1: Offline Indexing (Write Path)Phase 2: Online Querying (Read Path)
Execution TimingAsynchronous / Batch job (runs offline or on file upload).Synchronous / Real-time (runs while the user waits).
Input & OutputInput: Raw Documents ➔ Output: Populated Vector DB.Input: User Question ➔ Output: Grounded Final Answer.
Latency TargetSeconds to minutes (throughput matters more than speed).Under 1 second to first streamed token!
AI Models UsedEmbedding Model only (unless using LLM for chunk summaries).Both Embedding Model (for query) + LLM (for answer).
⚡ Knowledge Check

During Offline Indexing (when you load, chunk, and embed 500 PDFs into your Vector Database), does the Large Language Model (like GPT-4 or Llama-3) need to be called at all in a standard RAG setup?

A) No! Standard indexing only uses a fast, cheap Embedding Model to turn text chunks into vectors; the LLM is only called during live Querying▼
✓ Correct!Embedding models are 50x–100x cheaper and faster than generative LLMs. During indexing, only the Embedding Model runs to vectorize your chunks!
B) Yes, the generative LLM must rewrite every word of every PDF during indexing▼
✕ Incorrect.Standard RAG stores the original text chunks verbatim alongside their embedding vectors without calling an LLM during ingestion.

  1. Advanced: Naive RAG vs. Production Modular RAG (Rerankers & Incremental Sync)

As you move from a weekend prototype (Naive RAG) to an enterprise system (Production Modular RAG), engineers add three crucial upgrades:

Upgrade 1 · Pre-Retrieval
Query Rewriting

Users type messy, vague follow-up questions like "What about for contractors?" A fast LLM rewrites that into a standalone search query ("What is the leave policy for contractors?") before hitting the Vector DB!

Upgrade 2 · Post-Retrieval
Cross-Encoder Reranking

Retrieves the top 20 candidate chunks fast from the Vector DB, then uses a high-precision Reranker model to score and pick the true top 3 chunks before passing them to the LLM.

Upgrade 3 · Index Hygiene
Document Hash Syncing

Stores a content hash (SHA-256) for every file so when 1 page of a PDF is edited, the Indexing Job only re-embeds that 1 changed document and deletes stale chunks!

  1. Visual Explanation: The Complete 2-Phase RAG Blueprint

Trace the complete End-to-End RAG Blueprint below: watch how Phase 1 (Offline Indexing) populates the Vector Database at the center, and how Phase 2 (Online Querying) retrieves from that exact database to power the LLM!

🏗️️ PHASE 1 · OFFLINE INDEXING

Write Path

📄 1. Documents (PDFs, Docs, Wiki)

↓

🧹 2. Document Loading & Cleaning

↓

✂️ 3. Text Splitting & Chunking (500 tokens)

↓

🔢 4. Embedding Model ➔ Vectors

Shared Hub

🗄️ Vector DB

Stores Vectors + Text + Metadata

⚡ PHASE 2 · ONLINE QUERYING

Read Path

💬 1. User Query ➔ 🔢 Query Embedding

↓

🔍 2. Similarity Search (Fetches Top-K Chunks)

↓

📋 3. Augmented Prompt (Query + Chunks)

↓

🤖 4. LLM Generation ➔ Cited Answer!

⏱️ LATENCY BREAKDOWN INSPECTOR · CLICK TO EXPAND

Where Does the 650ms Go During a Live User Query?

Inspect

▼

1. Query Embedding
~35 ms
Converts question string into 1 vector
2. Vector DB Search
~15 ms
HNSW index finds top 3 chunks out of 1M
3. LLM First Token
~600 ms
Reads retrieved chunks & starts streaming

  1. Python Implementation: Object-Oriented 2-Phase RAG System

Here is a clean, runnable Python blueprint that models the exact separation between Phase 1 (indexDocument) and Phase 2 (answerUserQuery) using pure NumPy cosine similarity:

two_phase_rag_system.pyPython 3.11+ · NumPy Vector Index & Query Blueprint
import numpy as np # Shared Embedding Function (Used by BOTH Indexing and Querying!)def embedText(text: str) -> np.ndarray:    words = text.lower()    vec = np.array([        float("leave" in words or "vacation" in words or "days" in words),        float("vpn" in words or "password" in words or "reset" in words),        float("laptop" in words or "monitor" in words or "stipend" in words),    ], dtype=float)    norm = np.linalg.norm(vec)    return vec / norm if norm > 0 else vec class RagSystem:    def init(self):        self.vectorDb = []   # Stores (vector, chunkText, sourceMeta)     # PHASE 1: OFFLINE INDEXING (Load -> Chunk -> Embed -> Store)    def indexDocument(self, sourceName: str, fullText: str):        chunks = [c.strip() for c in fullText.split(".") if c.strip()]        for chunk in chunks:            vec = embedText(chunk)            self.vectorDb.append((vec, chunk, sourceName))     # PHASE 2: ONLINE QUERYING (Query Embed -> Search -> Augment Prompt)    def answerUserQuery(self, userQuery: str) -> str:        qVec = embedText(userQuery)        scored = [(float(np.dot(qVec, v)), txt, src) for v, txt, src in self.vectorDb]        bestScore, bestChunk, bestSource = max(scored, key=lambda item: item[0])        return "[Source: " + bestSource + " | Sim: " + str(round(bestScore, 2)) + "] -> " + bestChunk rag = RagSystem()rag.indexDocument("HR_Manual.pdf", "Employees get 22 vacation days. Unused leave carries over.")rag.indexDocument("IT_Guide.pdf",  "To reset your VPN password visit the IT portal. Monitors have a stipend.") print("Indexed Chunks in Vector DB:", len(rag.vectorDb))print(rag.answerUserQuery("How do I reset my VPN password?"))

Pro Tip (Look at How Fast answerUserQuery Runs!):

Notice that rag.indexDocument(...) did all the document splitting and chunk embedding ahead of time. When rag.answerUserQuery(...) is called, it only embeds the single user question and computes a fast dot product against the pre-computed vectors in self.vectorDb!

Key Points

✓Every RAG system is divided into two distinct workflows: the Offline Indexing Phase (Write Path) and the Online Query & Generation Phase (Read Path).
✓The Indexing Phase runs ahead of time: Documents ➔ Document Loading ➔ Text Splitting / Chunking ➔ Embedding Generation ➔ Vector Database Storage.
✓The Query Phase runs live per user question: User Query ➔ Query Embedding ➔ Vector Similarity Search ➔ Top-K Relevant Chunks ➔ Augmented Prompt ➔ LLM Final Answer.
✓Both phases must use the exact same Embedding Model so that query vectors and document chunk vectors live in the same mathematical vector space.
✓Production RAG architectures enhance this baseline with Query Rewriting (pre-retrieval), Cross-Encoder Reranking (post-retrieval), and Incremental Hash Syncing.

Common Mistakes

✕ Embedding the user query with a different model than the one used to index the documents.

If you switch from OpenAI's text-embedding-3-small to a HuggingFace BGE model, you must re-index your Vector Database—never compare vectors from two different models!

✕ Storing only the numeric vector in the Vector Database and forgetting to store the original chunk text and metadata!

An LLM cannot read a raw embedding vector! Your Vector Database must store the human-readable chunk_text and source_metadata alongside each vector so you can paste the text into the prompt.

✕ Re-indexing all 10,000 PDFs from scratch without checking if a document already exists in the Vector DB.

Calling vector_db.add() repeatedly on the same files creates duplicate chunks that crowd out diverse results during Top-K retrieval! Always use deterministic chunk IDs or upsert by file hash.

The Big Picture

Phase 1: Offline Indexing (Preparing the Knowledge Base)

Raw Documents → Load & Clean → Split into Chunks → Embed into Vectors → Store in Vector DB

Phase 2: Online Querying (Answering the User in Real Time)

User Query → Embed Query → Similarity Search in Vector DB → Top-K Chunks + Prompt → LLM Answer!

The big takeaway is simple: the Vector Database is the bridge where Offline Indexing meets Online Querying. Everything we do in Modules 4, 5, 6, and 7 builds that bridge, and everything in Modules 8, 9, and 10 crosses it!

Up Next in Module 4: Let's zoom into the very first station of Offline Indexing—Documents & Data Preparation (04-document-processing.mdx)—where we learn how to extract clean text and rich metadata from PDFs, Word docs, HTML web pages, Markdown, CSVs, and JSON!