The Core Idea: In RAG, there is an iron law: Garbage In, Garbage Out. Even if you use the smartest LLM on Earth and the fastest Vector Database, your RAG assistant will fail if you feed it broken PDF tables, leftover HTML navigation menus, or unlabelled text blobs! Document Processing & Data Preparation is the crucial first step where we load raw files (PDF, DOCX, HTML, Markdown, CSV, JSON), scrub away the noise, and wrap every piece of text inside a structured Document Object (Content + Metadata).
- Beginner: Where Does RAG Knowledge Come From? (The 7 File Formats)
In the real world, company knowledge rarely lives in clean plain-text files. Instead, it is scattered across seven common document formats—and each format has its own hidden traps when you extract text from it!
| File Format | Where You See It | Hidden Extraction Trap | Popular Python Loaders |
|---|---|---|---|
| 1. PDF (.pdf) | Manuals, Contracts, Research Papers | 2-column layouts & tables get scrambled left-to-right; scanned PDFs need OCR! | PyMuPDF (fitz), Docling, Llamaparse |
| 2. Word (.docx) | HR Policies, Proposals, Legal SOPs | Hidden XML tags, tracked changes, and nested tables. | python-docx, Unstructured |
| 3. HTML / Web | Help Centers, Blogs, Documentation Sites | 80% of raw HTML is boilerplate noise (navbars, cookie banners, footers, scripts)! | BeautifulSoup4, Trafilatura |
| 4. Markdown (.md) | GitHub Readmes, Notion Exports, Wikis | The Gold Standard for RAG! Preserves clean headings (#, ##) and tables. | Native Python / MarkdownHeaderSplitter |
| 5. CSV (.csv) | Product Catalogs, Pricing Sheets, Logs | A raw row like 104, 42.50, True makes zero sense without column names attached! | pandas, csv.DictReader |
| 6. JSON & TXT | API Responses, Slack/Ticket Exports | Deeply nested JSON keys need flattening into readable key-value sentences. | json, jq |
Why Modern RAG Engineers Convert PDFs & HTML Into Markdown First!
When a naive PDF reader extracts text from a pricing table, it strips all borders and dumps a soup of numbers onto one line. Modern document parsers (like PyMuPDF4LLM, Docling, and Marker) convert PDFs and HTML pages into clean Markdown—turning visual headings into ## Heading and tables into Markdown pipe tables (| Plan | Price |) that LLMs can read effortlessly!
- Beginner: The Anatomy of a RAG Document Object (Content + Metadata)
When a Document Loader reads a file, it does not just return a bare string of text. In every major RAG framework (LangChain, LlamaIndex, Haystack), it returns a standardized Document Object made of two parts:
The actual cleaned text extracted from the page or section (for example: "Full-time employees receive 22 days of paid annual leave per calendar year.").
A structured dictionary of facts about where that text came from:
• source: "HR_Handbook_2026.pdf"
• page: 14
• title: "Annual Leave Policy"
• category: "human_resources"
3 Reasons Why Metadata Is a Superpower in Later Modules
source and page in the metadata!category == "human_resources" first—skipping 90% of irrelevant engineering files!source == "HR_Handbook_2025.pdf" in one command.If you extract text from 500 PDFs and store the text chunks in your Vector Database WITHOUT attaching a metadata dictionary (no source filename or page number), what problem will happen when a user asks a question?
A) You cannot filter by document category/date during search, and the LLM will have no way to cite which PDF or page number the answer came from!▼
metadata dictionary during Document Loading!B) The Embedding Model automatically guesses the PDF filename from the text▼
page_content; they never know the file path or page number unless you store it in metadata.
- Medium: Text Cleaning & Normalization (Scrubbing Away the Noise)
Raw text extracted from PDFs and websites is full of digital trash that pollutes your embedding vectors and wastes LLM tokens. Before chunking any document, run it through this 4-Step Cleaning Checklist:
Every page of a corporate PDF often repeats "CONFIDENTIAL — ACME CORP 2026 — Page 4 of 90" right in the middle of a sentence that wraps across two pages! Strip repeated header/footer banners so sentences flow cleanly.
When scraping a webpage, strip <nav>, <footer>, <script>, <style>, and cookie consent banners. Otherwise, every webpage in your database will have identical vectors for "Home | About | Privacy Policy"!
PDFs often split words across lines like "reim-\nbursement" and insert 15 blank newlines between paragraphs. Rejoin hyphenated line-wraps and collapse multiple spaces/newlines into clean single breaks.
PDF fonts often encode ligatures like "fi" (in "finance") as a single weird symbol instead of "f" + "i", or insert zero-width spaces (\u200b). Running unicodedata.normalize("NFKC", text) converts them back to standard characters!
- Medium: Turning Tabular CSV & JSON Data Into Self-Contained Text
What happens if you load a CSV spreadsheet of products and embed each row as raw comma-separated values? Look at the difference between a Bad Raw Row and a Serialized Key-Value Row:
"ProBook-15, 1299, 16, 512, True, Mumbai"
What is 16? What is 512? What is True? Neither the Embedding Model nor the LLM knows what those numbers mean because the column headers were left at the top of the file!
"Product: ProBook-15 | Price: USD 1299 | RAM: 16GB | Storage: 512GB SSD | In Stock: True | Warehouse: Mumbai"
By gluing the column header to every cell value, every single row becomes a 100% self-contained, human-readable sentence!
When preparing a 1,000-row CSV spreadsheet for a RAG system, why should you attach the column names to each row's text (e.g., "Model: X1 | Price: USD 99") instead of just splitting the file by lines?
A) Because Row #500 gets retrieved by itself without Row #1 (the header row), so each row needs its column labels attached to have meaning!▼
B) Because Vector Databases cannot store numbers unless they are written in uppercase▼
- Advanced: Handling Scanned PDFs (OCR), Complex Tables & PII Redaction
In enterprise RAG, three advanced data-preparation challenges separate junior prototypes from production systems:
Some PDFs are just scanned photos of paper—standard loaders return an empty string ""! Production loaders check if len(extracted_text) < 50 on a page and automatically route that page through OCR / Vision-Language Models (Tesseract, Surya, or GPT-4o Vision).
Financial reports and datasheets pack critical facts inside multi-column tables. Using layout-aware parsers (Docling, Unstructured, or Llamaparse) detects table bounding boxes and converts them into clean Markdown or HTML tables.
Before indexing customer support tickets or medical notes, run a PII Redactor (like Microsoft Presidio or regex) to mask credit card numbers, SSNs, and phone numbers into [REDACTED_PHONE]!
- Visual Explanation: The Document Scrubber & Metadata Packaging Studio
See what happens inside a RAG Document Processor below: compare a Raw Noisy Extraction against a Cleaned & Normalized Document Object complete with its structured Metadata Tree!
❌ BEFORE · RAW UNCLEANED EXTRACTION
Polluted Signal
Sec-tion 3.2: Full-time employ-\nees receive 22 days of paid\n\n\nannual leave. Contact john.doe@acme.com or 555-019-2834.
⚠️ Contains HTML navbars, page headers, broken hyphenated words (employ-\nees), extra newlines, and unredacted phone numbers!
✅ AFTER · STRUCTURED DOCUMENT OBJECT
Ready for Chunking
├── content: "Section 3.2: Full-time employees receive 22 days of paid annual leave. Contact [REDACTED_EMAIL] or [REDACTED_PHONE]."
└── metadata:
├── source: "HR_Handbook_2026.pdf"
├── page: 12
├── title: "Section 3.2: Annual Leave"
└── category: "human_resources"
★ 100% clean prose + rich searchable metadata attached!
- Python Implementation: Building a Text Cleaner & Document Object Builder
Here is a complete, runnable Python script that implements a production-grade RAG Text Cleaner (Unicode normalization, HTML tag removal, broken hyphen repair, whitespace collapsing, and CSV row contextualization) and packages the result into a structured Document Object:
import reimport unicodedata # 1. Standardized RAG Document Container (Content + Metadata)class RagDocument: def init(self, content: str, source: str, page: int, title: str, category: str): self.content = content self.metadata = dict(source=source, page=page, title=title, category=category) # 2. 4-Step Text Cleaning & Normalization Functiondef cleanDocumentText(rawText: str) -> str: # Step A: Normalize Unicode ligatures (e.g., 'fi' -> 'fi') text = unicodedata.normalize("NFKC", rawText) # Step B: Strip HTML tags and repeated CONFIDENTIAL page headers text = re.sub(r"<[^>]+>", " ", text) text = re.sub(r"CONFIDENTIAL\s*-\sACME CORP\s-\sPAGE\s\d+", " ", text) # Step C: Rejoin words split across PDF line breaks (e.g., 'employ-\nees' -> 'employees') text = re.sub(r"(\w+)-\s*\n\s*(\w+)", r"\1\2", text) # Step D: Collapse multiple spaces and newlines into a single space return re.sub(r"\s+", " ", text).strip() noisyInput = "<nav>Menu</nav> CONFIDENTIAL - ACME CORP - PAGE 12 Full-time employ-\nees get 22 days of financial leave."cleanText = cleanDocumentText(noisyInput)docObject = RagDocument(cleanText, source="HR_Handbook.pdf", page=12, title="Annual Leave", category="hr") print("Cleaned Content:", docObject.content)print("Attached Metadata:", docObject.metadata)Pro Tip (Look at Step C in cleanDocumentText Above!):
The regular expression re.sub(r"(\w+)-\s*\n\s*(\w+)", r"\1\2", text) fixes one of the most annoying bugs in PDF RAG systems: when a long keyword like "reimbursement" gets hyphenated at the right margin of a PDF as "reim-\nbursement", tokenizers split it into nonsense subwords unless you stitch it back together first!
Key Points
content (the cleaned text) with metadata (source, page, title, category).Common Mistakes
✕ Using a basic text-only PDF loader on scanned image PDFs without checking if extracted text is empty.
Standard PDF loaders only read digital text layers; on a scanned contract, they silently return 0 characters! Always check text length per page and trigger OCR when a page is blank.
✕ Indexing raw HTML web pages without stripping navigation menus, sidebars, and footers.
Leaving website menus and footers in your text creates hundreds of chunks filled with links instead of actual help-center content.
✕ Discarding metadata during document loading or text splitting.
When you split a 10-page document into 40 chunks in the next step, make sure every single child chunk inherits its parent document's metadata dictionary!
The Big Picture
Raw Unstructured Files (Messy Inputs)
PDFs + Word Docs + HTML Pages + CSVs → Full of Headers, Broken Hyphens & Boilerplate Noise
Document Processing Station (Clean Outputs)
Parse Layout → Scrub & Normalize Text → Attach Metadata (Source, Page, Category) → Ready to Chunk!
The big takeaway is simple: clean data beats clever algorithms every single time. By scrubbing away formatting noise and packaging every page with rich metadata right at the doorway, you set the rest of your RAG system up for high-precision retrieval.
Up Next in Module 5: Now that we have clean, metadata-tagged Document objects, how do we slice a 50-page document into bite-sized pieces without cutting sentences or tables in half? Let's master Chunking & Text Splitting (05-chunking.mdx)!