The Core Idea: A beginner often thinks a RAG system reads and processes 1,000 PDFs from scratch every single time a user asks a question. If it did that, every answer would take 20 minutes! Instead, every production RAG system is split into two separate workflows: (1) the Offline Indexing Phase (which prepares, chops, embeds, and organizes your documents into a Vector Database ahead of time—like stocking a library's shelves), and (2) the Online Query Phase (which searches that pre-built index in milliseconds to answer a live user question).
- Beginner: The Two-Phase Secret (Stocking the Library vs. Answering a Visitor)
Imagine walking into a city library and asking the librarian: "Which book explains how solar panels work?"
The librarian does not open all 50,000 books on the floor and read them while you wait! Months before you arrived, the library ran an Indexing Process—cataloging every chapter by topic so the librarian can look up the exact shelf in 5 seconds.
A RAG System works the exact same way by separating its work into two phases:
- Beginner: Step-by-Step Through Phase 1 (Offline Document Indexing)
How do we turn a messy folder of company PDFs, Word docs, and web pages into a searchable AI memory bank? We pass the documents through five sequential stations during Offline Indexing:
Raw files: PDFs, DOCX, Markdown, HTML web pages, Notion wikis, or CSVs.
Extracts clean text strings and attaches metadata (filename, page number, author).
Slices long 50-page documents into bite-sized passages (e.g., 300 to 500 words each).
Converts each text chunk into a dense meaning vector (e.g., 1,536 numbers).
Stores the vectors + original text + metadata together in a fast search index!
Why do we split a 100-page PDF into smaller chunks (Stage 3) before creating embeddings, instead of turning the entire 100-page book into one single vector?
A) One vector for 100 pages blurs all topics together, and we only want to pass the specific relevant paragraph to the LLM▼
B) Because hard drives cannot store files larger than 1 page▼
- Medium: Step-by-Step Through Phase 2 (Online Retrieval & Generation)
Once the Vector Database is stocked and ready, what happens in the 500 milliseconds after a user types a question into your app? The Online Query Phase executes 5 live steps:
"How do I reset my corporate VPN password?"
Query Vector q = [0.14, -0.62, 0.88, ...]
Cosine Similarity ➔ Top-K (e.g., Top 3 Chunks)
Grounded Prompt Template
"Go to vpn.company.com and click..."
⚠️ The #1 Golden Rule of RAG: Never Mix Embedding Models!
Notice Live Step 2 above: the model that embeds the User Query in Phase 2 MUST be the exact same Embedding Model that embedded the Document Chunks in Phase 1! Different embedding models speak completely different mathematical coordinate languages—comparing a query vector from Model A to document vectors from Model B produces random garbage!
- Medium: Offline Indexing vs. Online Querying — Head-to-Head Comparison
Understanding the separation between Offline Indexing and Online Querying is essential for designing and debugging any RAG system:
| System Aspect | Phase 1: Offline Indexing (Write Path) | Phase 2: Online Querying (Read Path) |
|---|---|---|
| Execution Timing | Asynchronous / Batch job (runs offline or on file upload). | Synchronous / Real-time (runs while the user waits). |
| Input & Output | Input: Raw Documents ➔ Output: Populated Vector DB. | Input: User Question ➔ Output: Grounded Final Answer. |
| Latency Target | Seconds to minutes (throughput matters more than speed). | Under 1 second to first streamed token! |
| AI Models Used | Embedding Model only (unless using LLM for chunk summaries). | Both Embedding Model (for query) + LLM (for answer). |
During Offline Indexing (when you load, chunk, and embed 500 PDFs into your Vector Database), does the Large Language Model (like GPT-4 or Llama-3) need to be called at all in a standard RAG setup?
A) No! Standard indexing only uses a fast, cheap Embedding Model to turn text chunks into vectors; the LLM is only called during live Querying▼
B) Yes, the generative LLM must rewrite every word of every PDF during indexing▼
- Advanced: Naive RAG vs. Production Modular RAG (Rerankers & Incremental Sync)
As you move from a weekend prototype (Naive RAG) to an enterprise system (Production Modular RAG), engineers add three crucial upgrades:
Users type messy, vague follow-up questions like "What about for contractors?" A fast LLM rewrites that into a standalone search query ("What is the leave policy for contractors?") before hitting the Vector DB!
Retrieves the top 20 candidate chunks fast from the Vector DB, then uses a high-precision Reranker model to score and pick the true top 3 chunks before passing them to the LLM.
Stores a content hash (SHA-256) for every file so when 1 page of a PDF is edited, the Indexing Job only re-embeds that 1 changed document and deletes stale chunks!
- Visual Explanation: The Complete 2-Phase RAG Blueprint
Trace the complete End-to-End RAG Blueprint below: watch how Phase 1 (Offline Indexing) populates the Vector Database at the center, and how Phase 2 (Online Querying) retrieves from that exact database to power the LLM!
🏗️️ PHASE 1 · OFFLINE INDEXING
Write Path
📄 1. Documents (PDFs, Docs, Wiki)
🧹 2. Document Loading & Cleaning
✂️ 3. Text Splitting & Chunking (500 tokens)
🔢 4. Embedding Model ➔ Vectors
Shared Hub
🗄️ Vector DB
Stores Vectors + Text + Metadata
⚡ PHASE 2 · ONLINE QUERYING
Read Path
💬 1. User Query ➔ 🔢 Query Embedding
🔍 2. Similarity Search (Fetches Top-K Chunks)
📋 3. Augmented Prompt (Query + Chunks)
🤖 4. LLM Generation ➔ Cited Answer!
⏱️ LATENCY BREAKDOWN INSPECTOR · CLICK TO EXPAND
Where Does the 650ms Go During a Live User Query?
Inspect▼
⏱️ LATENCY BREAKDOWN INSPECTOR · CLICK TO EXPAND
Where Does the 650ms Go During a Live User Query?
▼
- Python Implementation: Object-Oriented 2-Phase RAG System
Here is a clean, runnable Python blueprint that models the exact separation between Phase 1 (indexDocument) and Phase 2 (answerUserQuery) using pure NumPy cosine similarity:
import numpy as np # Shared Embedding Function (Used by BOTH Indexing and Querying!)def embedText(text: str) -> np.ndarray: words = text.lower() vec = np.array([ float("leave" in words or "vacation" in words or "days" in words), float("vpn" in words or "password" in words or "reset" in words), float("laptop" in words or "monitor" in words or "stipend" in words), ], dtype=float) norm = np.linalg.norm(vec) return vec / norm if norm > 0 else vec class RagSystem: def init(self): self.vectorDb = [] # Stores (vector, chunkText, sourceMeta) # PHASE 1: OFFLINE INDEXING (Load -> Chunk -> Embed -> Store) def indexDocument(self, sourceName: str, fullText: str): chunks = [c.strip() for c in fullText.split(".") if c.strip()] for chunk in chunks: vec = embedText(chunk) self.vectorDb.append((vec, chunk, sourceName)) # PHASE 2: ONLINE QUERYING (Query Embed -> Search -> Augment Prompt) def answerUserQuery(self, userQuery: str) -> str: qVec = embedText(userQuery) scored = [(float(np.dot(qVec, v)), txt, src) for v, txt, src in self.vectorDb] bestScore, bestChunk, bestSource = max(scored, key=lambda item: item[0]) return "[Source: " + bestSource + " | Sim: " + str(round(bestScore, 2)) + "] -> " + bestChunk rag = RagSystem()rag.indexDocument("HR_Manual.pdf", "Employees get 22 vacation days. Unused leave carries over.")rag.indexDocument("IT_Guide.pdf", "To reset your VPN password visit the IT portal. Monitors have a stipend.") print("Indexed Chunks in Vector DB:", len(rag.vectorDb))print(rag.answerUserQuery("How do I reset my VPN password?"))Pro Tip (Look at How Fast answerUserQuery Runs!):
Notice that rag.indexDocument(...) did all the document splitting and chunk embedding ahead of time. When rag.answerUserQuery(...) is called, it only embeds the single user question and computes a fast dot product against the pre-computed vectors in self.vectorDb!
Key Points
Common Mistakes
✕ Embedding the user query with a different model than the one used to index the documents.
If you switch from OpenAI's text-embedding-3-small to a HuggingFace BGE model, you must re-index your Vector Database—never compare vectors from two different models!
✕ Storing only the numeric vector in the Vector Database and forgetting to store the original chunk text and metadata!
An LLM cannot read a raw embedding vector! Your Vector Database must store the human-readable chunk_text and source_metadata alongside each vector so you can paste the text into the prompt.
✕ Re-indexing all 10,000 PDFs from scratch without checking if a document already exists in the Vector DB.
Calling vector_db.add() repeatedly on the same files creates duplicate chunks that crowd out diverse results during Top-K retrieval! Always use deterministic chunk IDs or upsert by file hash.
The Big Picture
Phase 1: Offline Indexing (Preparing the Knowledge Base)
Raw Documents → Load & Clean → Split into Chunks → Embed into Vectors → Store in Vector DB
Phase 2: Online Querying (Answering the User in Real Time)
User Query → Embed Query → Similarity Search in Vector DB → Top-K Chunks + Prompt → LLM Answer!
The big takeaway is simple: the Vector Database is the bridge where Offline Indexing meets Online Querying. Everything we do in Modules 4, 5, 6, and 7 builds that bridge, and everything in Modules 8, 9, and 10 crosses it!
Up Next in Module 4: Let's zoom into the very first station of Offline Indexing—Documents & Data Preparation (04-document-processing.mdx)—where we learn how to extract clean text and rich metadata from PDFs, Word docs, HTML web pages, Markdown, CSVs, and JSON!