Documents & Data Preparation: Loading, Cleaning & Metadata for RAG

Learn where a RAG system's knowledge actually comes from—how to parse PDFs, DOCX, HTML, Markdown, CSV, and JSON, clean noisy text, and attach rich Metadata for precision retrieval.

28 minBeginnerCode Examples

The Core Idea: In RAG, there is an iron law: Garbage In, Garbage Out. Even if you use the smartest LLM on Earth and the fastest Vector Database, your RAG assistant will fail if you feed it broken PDF tables, leftover HTML navigation menus, or unlabelled text blobs! Document Processing & Data Preparation is the crucial first step where we load raw files (PDF, DOCX, HTML, Markdown, CSV, JSON), scrub away the noise, and wrap every piece of text inside a structured Document Object (Content + Metadata).

  1. Beginner: Where Does RAG Knowledge Come From? (The 7 File Formats)

In the real world, company knowledge rarely lives in clean plain-text files. Instead, it is scattered across seven common document formats—and each format has its own hidden traps when you extract text from it!

File FormatWhere You See ItHidden Extraction TrapPopular Python Loaders
1. PDF (.pdf)Manuals, Contracts, Research Papers2-column layouts & tables get scrambled left-to-right; scanned PDFs need OCR!PyMuPDF (fitz), Docling, Llamaparse
2. Word (.docx)HR Policies, Proposals, Legal SOPsHidden XML tags, tracked changes, and nested tables.python-docx, Unstructured
3. HTML / WebHelp Centers, Blogs, Documentation Sites80% of raw HTML is boilerplate noise (navbars, cookie banners, footers, scripts)!BeautifulSoup4, Trafilatura
4. Markdown (.md)GitHub Readmes, Notion Exports, WikisThe Gold Standard for RAG! Preserves clean headings (#, ##) and tables.Native Python / MarkdownHeaderSplitter
5. CSV (.csv)Product Catalogs, Pricing Sheets, LogsA raw row like 104, 42.50, True makes zero sense without column names attached!pandas, csv.DictReader
6. JSON & TXTAPI Responses, Slack/Ticket ExportsDeeply nested JSON keys need flattening into readable key-value sentences.json, jq

Why Modern RAG Engineers Convert PDFs & HTML Into Markdown First!

When a naive PDF reader extracts text from a pricing table, it strips all borders and dumps a soup of numbers onto one line. Modern document parsers (like PyMuPDF4LLM, Docling, and Marker) convert PDFs and HTML pages into clean Markdown—turning visual headings into ## Heading and tables into Markdown pipe tables (| Plan | Price |) that LLMs can read effortlessly!

  1. Beginner: The Anatomy of a RAG Document Object (Content + Metadata)

When a Document Loader reads a file, it does not just return a bare string of text. In every major RAG framework (LangChain, LlamaIndex, Haystack), it returns a standardized Document Object made of two parts:

1. page_content (The Text String)What Gets Embedded

The actual cleaned text extracted from the page or section (for example: "Full-time employees receive 22 days of paid annual leave per calendar year.").

2. metadata (The Key-Value Dictionary)The ID Tag & Filter

A structured dictionary of facts about where that text came from:
• source: "HR_Handbook_2026.pdf"
• page: 14
• title: "Annual Leave Policy"
• category: "human_resources"

3 Reasons Why Metadata Is a Superpower in Later Modules

1. Verifiable Citations (Module 9): The LLM can only say "According to HR_Handbook_2026.pdf, Page 14..." if you saved source and page in the metadata!
2. Pre-Filtering Search (Module 7 & 8): If a user asks about HR policies, your Vector DB can filter category == "human_resources" first—skipping 90% of irrelevant engineering files!
3. Deleting Old Versions: When a 2025 PDF is replaced by a 2026 PDF, you can delete all chunks where source == "HR_Handbook_2025.pdf" in one command.
⚡ Knowledge Check

If you extract text from 500 PDFs and store the text chunks in your Vector Database WITHOUT attaching a metadata dictionary (no source filename or page number), what problem will happen when a user asks a question?

A) You cannot filter by document category/date during search, and the LLM will have no way to cite which PDF or page number the answer came from!▼
✓ Correct!A bare text chunk has no memory of which file or page it came from unless you attach that information in its metadata dictionary during Document Loading!
B) The Embedding Model automatically guesses the PDF filename from the text▼
✕ Incorrect.Embedding models only see the characters inside page_content; they never know the file path or page number unless you store it in metadata.

  1. Medium: Text Cleaning & Normalization (Scrubbing Away the Noise)

Raw text extracted from PDFs and websites is full of digital trash that pollutes your embedding vectors and wastes LLM tokens. Before chunking any document, run it through this 4-Step Cleaning Checklist:

1. Strip Repeated Headers, Footers & Page Numbers

Every page of a corporate PDF often repeats "CONFIDENTIAL — ACME CORP 2026 — Page 4 of 90" right in the middle of a sentence that wraps across two pages! Strip repeated header/footer banners so sentences flow cleanly.

2. Remove HTML Boilerplate (Navbars & Footers)

When scraping a webpage, strip <nav>, <footer>, <script>, <style>, and cookie consent banners. Otherwise, every webpage in your database will have identical vectors for "Home | About | Privacy Policy"!

3. Fix Broken Line-Wrap Hyphens & Whitespace

PDFs often split words across lines like "reim-\nbursement" and insert 15 blank newlines between paragraphs. Rejoin hyphenated line-wraps and collapse multiple spaces/newlines into clean single breaks.

4. Unicode Normalization (NFKC)

PDF fonts often encode ligatures like "fi" (in "finance") as a single weird symbol instead of "f" + "i", or insert zero-width spaces (\u200b). Running unicodedata.normalize("NFKC", text) converts them back to standard characters!

  1. Medium: Turning Tabular CSV & JSON Data Into Self-Contained Text

What happens if you load a CSV spreadsheet of products and embed each row as raw comma-separated values? Look at the difference between a Bad Raw Row and a Serialized Key-Value Row:

❌ Bad: Raw CSV Row Without Headers

"ProBook-15, 1299, 16, 512, True, Mumbai"

What is 16? What is 512? What is True? Neither the Embedding Model nor the LLM knows what those numbers mean because the column headers were left at the top of the file!

✅ Good: Header-Contextualized Row

"Product: ProBook-15 | Price: USD 1299 | RAM: 16GB | Storage: 512GB SSD | In Stock: True | Warehouse: Mumbai"

By gluing the column header to every cell value, every single row becomes a 100% self-contained, human-readable sentence!

⚡ Knowledge Check

When preparing a 1,000-row CSV spreadsheet for a RAG system, why should you attach the column names to each row's text (e.g., "Model: X1 | Price: USD 99") instead of just splitting the file by lines?

A) Because Row #500 gets retrieved by itself without Row #1 (the header row), so each row needs its column labels attached to have meaning!▼
✓ Correct!When Row #500 is retrieved from a Vector Database, it is separated from the top of the spreadsheet. Attaching the column names to every row ensures both the embedding model and LLM understand every value!
B) Because Vector Databases cannot store numbers unless they are written in uppercase▼
✕ Incorrect.Vector databases store any text string; the issue is preserving semantic context for each individual row.

  1. Advanced: Handling Scanned PDFs (OCR), Complex Tables & PII Redaction

In enterprise RAG, three advanced data-preparation challenges separate junior prototypes from production systems:

1. Scanned PDFs & OCR
Detecting Image-Only PDFs

Some PDFs are just scanned photos of paper—standard loaders return an empty string ""! Production loaders check if len(extracted_text) < 50 on a page and automatically route that page through OCR / Vision-Language Models (Tesseract, Surya, or GPT-4o Vision).

2. Table Preservation
Layout-Aware Parsing

Financial reports and datasheets pack critical facts inside multi-column tables. Using layout-aware parsers (Docling, Unstructured, or Llamaparse) detects table bounding boxes and converts them into clean Markdown or HTML tables.

3. PII Scrubbing
Redacting Sensitive Data

Before indexing customer support tickets or medical notes, run a PII Redactor (like Microsoft Presidio or regex) to mask credit card numbers, SSNs, and phone numbers into [REDACTED_PHONE]!

  1. Visual Explanation: The Document Scrubber & Metadata Packaging Studio

See what happens inside a RAG Document Processor below: compare a Raw Noisy Extraction against a Cleaned & Normalized Document Object complete with its structured Metadata Tree!

❌ BEFORE · RAW UNCLEANED EXTRACTION

Polluted Signal

<nav>Home | Careers | Login | Cookie Policy</nav>
CONFIDENTIAL - ACME CORP - PAGE 12 OF 85

Sec-tion 3.2: Full-time employ-\nees receive 22 days of paid\n\n\nannual leave. Contact john.doe@acme.com or 555-019-2834.

<footer>Copyright 2026 All Rights Reserved</footer>

⚠️ Contains HTML navbars, page headers, broken hyphenated words (employ-\nees), extra newlines, and unredacted phone numbers!

✅ AFTER · STRUCTURED DOCUMENT OBJECT

Ready for Chunking

Document

├── content: "Section 3.2: Full-time employees receive 22 days of paid annual leave. Contact [REDACTED_EMAIL] or [REDACTED_PHONE]."

└── metadata:

├── source: "HR_Handbook_2026.pdf"
├── page: 12
├── title: "Section 3.2: Annual Leave"
└── category: "human_resources"

★ 100% clean prose + rich searchable metadata attached!

  1. Python Implementation: Building a Text Cleaner & Document Object Builder

Here is a complete, runnable Python script that implements a production-grade RAG Text Cleaner (Unicode normalization, HTML tag removal, broken hyphen repair, whitespace collapsing, and CSV row contextualization) and packages the result into a structured Document Object:

rag_document_processor.pyPython 3.11+ · Standard Library (re, unicodedata)
import reimport unicodedata # 1. Standardized RAG Document Container (Content + Metadata)class RagDocument:    def init(self, content: str, source: str, page: int, title: str, category: str):        self.content = content        self.metadata = dict(source=source, page=page, title=title, category=category) # 2. 4-Step Text Cleaning & Normalization Functiondef cleanDocumentText(rawText: str) -> str:    # Step A: Normalize Unicode ligatures (e.g., 'fi' -> 'fi')    text = unicodedata.normalize("NFKC", rawText)     # Step B: Strip HTML tags and repeated CONFIDENTIAL page headers    text = re.sub(r"<[^>]+>", " ", text)    text = re.sub(r"CONFIDENTIAL\s*-\sACME CORP\s-\sPAGE\s\d+", " ", text)     # Step C: Rejoin words split across PDF line breaks (e.g., 'employ-\nees' -> 'employees')    text = re.sub(r"(\w+)-\s*\n\s*(\w+)", r"\1\2", text)     # Step D: Collapse multiple spaces and newlines into a single space    return re.sub(r"\s+", " ", text).strip() noisyInput = "<nav>Menu</nav> CONFIDENTIAL - ACME CORP - PAGE 12 Full-time employ-\nees get   22   days of financial leave."cleanText  = cleanDocumentText(noisyInput)docObject  = RagDocument(cleanText, source="HR_Handbook.pdf", page=12, title="Annual Leave", category="hr") print("Cleaned Content:", docObject.content)print("Attached Metadata:", docObject.metadata)

Pro Tip (Look at Step C in cleanDocumentText Above!):

The regular expression re.sub(r"(\w+)-\s*\n\s*(\w+)", r"\1\2", text) fixes one of the most annoying bugs in PDF RAG systems: when a long keyword like "reimbursement" gets hyphenated at the right margin of a PDF as "reim-\nbursement", tokenizers split it into nonsense subwords unless you stitch it back together first!

Key Points

✓RAG systems ingest knowledge from diverse formats (PDF, DOCX, HTML, Markdown, CSV, JSON, TXT), and converting complex documents into clean Markdown preserves headings and tables best for LLMs.
✓Every loaded document is represented as a Document Object pairing content (the cleaned text) with metadata (source, page, title, category).
✓Metadata is essential for generating verifiable source citations, pre-filtering searches by category or date, and deleting outdated files from the Vector Database.
✓Text Cleaning removes HTML navigation clutter, repeated PDF headers/footers, broken line-wrap hyphens, extra whitespace, and weird Unicode ligatures before chunking.
✓When loading tabular CSV or JSON data, always attach column names to every cell value so each row remains self-contained after retrieval.

Common Mistakes

✕ Using a basic text-only PDF loader on scanned image PDFs without checking if extracted text is empty.

Standard PDF loaders only read digital text layers; on a scanned contract, they silently return 0 characters! Always check text length per page and trigger OCR when a page is blank.

✕ Indexing raw HTML web pages without stripping navigation menus, sidebars, and footers.

Leaving website menus and footers in your text creates hundreds of chunks filled with links instead of actual help-center content.

✕ Discarding metadata during document loading or text splitting.

When you split a 10-page document into 40 chunks in the next step, make sure every single child chunk inherits its parent document's metadata dictionary!

The Big Picture

Raw Unstructured Files (Messy Inputs)

PDFs + Word Docs + HTML Pages + CSVs → Full of Headers, Broken Hyphens & Boilerplate Noise

Document Processing Station (Clean Outputs)

Parse Layout → Scrub & Normalize Text → Attach Metadata (Source, Page, Category) → Ready to Chunk!

The big takeaway is simple: clean data beats clever algorithms every single time. By scrubbing away formatting noise and packaging every page with rich metadata right at the doorway, you set the rest of your RAG system up for high-precision retrieval.

Up Next in Module 5: Now that we have clean, metadata-tagged Document objects, how do we slice a 50-page document into bite-sized pieces without cutting sentences or tables in half? Let's master Chunking & Text Splitting (05-chunking.mdx)!