Python Syntax, Functions, and Data Structures

Master the essential Python syntax, memory-efficient data structures, comprehensions, and functional patterns that power modern Machine Learning and AI pipelines.

24 minBeginnerCode Examples

The Core Thesis: Python is the undisputed lingua franca of Artificial Intelligence not because it is the fastest language on a CPU, but because its expressive syntax acts as a high-level control layer over C++ and CUDA tensor engines. Mastering how Python stores data in Lists, Tuples, Dictionaries, and Sets—and how it routes logic through functions and generators—is the foundation for writing clean training loops, tokenizers, and data loaders.

  1. Python Syntax Essentials and Dynamic Typing

Unlike C++ or Java, Python uses Dynamic Typing (types are attached to objects in memory at runtime, not variable names) and enforces code blocks using 4-space indentation rather than curly braces. In modern AI engineering, however, we combine dynamic execution with Type Hints so IDEs and teams always know whether a variable holds a float, a list of tokens, or a tensor shape:

1. Variables & Types
int, float, str, bool

Variables are references (pointers) to objects in memory. Rebinding a variable points the name to a new object.

2. Control Flow
if / elif / else, for, while

Orchestrates training epochs, early stopping checks, and batch iteration using helpers like enumerate() and zip().

3. Reference vs. Value
is vs. ==

a == b checks if two objects have equal values, while a is b checks if both point to the exact same memory address.

AI Engineering Rule (Unpacking): Python allows Sequence Unpacking in a single line: batch_size, seq_len, hidden_dim = tensor_shape or for step, (inputs, labels) in enumerate(dataloader):. You will see this pattern in every single PyTorch script!

  1. The Four Core Built-in Data Structures

Every dataset, JSON API payload, vocabulary lookup table, and model configuration in Python is built from four core collection types. Choosing the wrong one can slow down a preprocessing script by 1,000×1{,}000\times:

1. List — [10, 20, 30]Mutable · Ordered

Dynamic array of references. Supports fast O(1)\mathcal{O}(1) indexing (x[0]) and appending (x.append()), plus slicing (x[start:stop:step]).

Used for: Collecting epoch losses, token sequences, and batch buffers.

2. Tuple — (64, 3, 224, 224)Immutable · Ordered

Fixed, read-only sequence that cannot be modified after creation. Uses less memory than a list and can be used as a dictionary key.

Used for: Tensor shapes, (input, label) pairs, and multi-value function returns.

3. Dictionary — {"lr": 0.001}Hash Map · O(1) Lookup

Maps unique hashable keys to values using a hash table. Looking up a key takes O(1)\mathcal{O}(1) constant time even with 100,000100{,}000 entries!

Used for: LLM Token-to-ID vocabularies, model state_dict weights, and JSON configs.

4. Set — {"cat", "dog"}Unique · O(1) Membership

Unordered collection of unique hashable elements. Automatically strips duplicates and computes item in my_set in O(1)\mathcal{O}(1) time.

Used for: Deduplicating training corpora, stop-word filtering, and vocabulary extraction.

Data StructureSyntaxMutable?Lookup Time (x in col)Primary AI Use Case
List[a, b, c]YesO(N)\mathcal{O}(N) Linear ScanOrdered sequences of samples, tokens, or layer modules
Tuple(a, b, c)No (Fixed)O(N)\mathcal{O}(N) Linear ScanTensor shapes (B, T, D) and immutable dataset records
Dict{k: v}YesO(1)\mathcal{O}(1) Instant HashTokenizer vocabularies (word -> id) and hyperparameters
Set{a, b, c}YesO(1)\mathcal{O}(1) Instant HashFast deduplication and checking train/test ID overlap (leakage)
⚡ Knowledge Check

You are filtering 10,000,00010{,}000{,}000 words in a text corpus against a list of 50,00050{,}000 banned stop-words using if word not in stop_words:. Why does storing stop_words as a Python set run thousands of times faster than storing it as a list?

A) Sets use a hash table for O(1) lookup; Lists scan O(N) elements every time▼
✓ Correct!Checking word in list checks up to 50,00050{,}000 items per word (O(N)\mathcal{O}(N)), totaling 5×10115 \times 10^{11} operations! A set hashes the string to jump directly to its memory bucket in O(1)\mathcal{O}(1) constant time.
B) Sets automatically run on the GPU▼
✕ Incorrect.Built-in Python sets run entirely on the CPU in RAM. Their massive speedup comes purely from O(1)\mathcal{O}(1) hash table indexing versus O(N)\mathcal{O}(N) linear list scanning.

  1. Indexing, Slicing, and Mutability Traps

Python sequences (Lists, Tuples, Strings, and later NumPy/PyTorch tensors) share a universal Zero-Based Indexing and Slicing syntax: seq[start : stop : step], where start is inclusive and stop is exclusive:

tokens[0:3]
Grabs indices 0,1,20, 1, 2 (first 33 items). Stop index 33 is excluded.
tokens[-1]
Negative index counts from the end (grabs the very last element).
tokens[::-1]
Step of −1-1 reverses the entire sequence cleanly.

The #1 Python Bug in ML: Reference Aliasing of Mutable Lists

When you write list_b = list_a, Python does NOT copy the list! Both variables point to the exact same list in memory. Mutating list_b.append(99) silently corrupts list_a! Always use list_a.copy() (shallow copy) or copy.deepcopy(list_a) (for nested 2D matrices/dicts) when duplicating mutable data.

  1. Comprehensions and Generators (Lazy Memory Pipelines)

Instead of writing verbose 4-line for loops with .append(), Python provides Comprehensions to transform and filter data in a single readable expression—and Generators to stream massive datasets without running out of RAM:

1. List & Dict Comprehensions (Eager)

Constructs the entire new list or dictionary in memory immediately. Ideal for normalizing feature arrays or inverting a tokenizer vocabulary:

# List: [expr for x in iterable if cond]squares = [x**2 for x in range(10) if x % 2 == 0]# Dict: Swap word->id to id->wordid2word = {idx: w for w, idx in vocab.items()}
2. Generators & yield (Lazy Evaluation)

Uses parentheses (expr for x in iterable) or the yield keyword inside a function. Produces one item at a time on demand using O(1)\mathcal{O}(1) memory!

# Streams 100GB file 1 batch at a timedef batchGenerator(data, size): for i in range(0, len(data), size): yield data[i : i + size]
⚡ Knowledge Check

You need to stream 50,000,00050{,}000{,}000 high-resolution images from disk into a neural network training loop on a machine with only 16 GB16\text{ GB} of RAM. Should your data loader use a List Comprehension [...] or a Generator yield?

A) A Generator (yield), because it loads one batch lazily in O(1) RAM▼
✓ Correct!A generator pauses after yielding each mini-batch and discards it once processed, keeping RAM usage tiny regardless of dataset size.
B) A List Comprehension [...], because lists compress images▼
✕ Incorrect.A list comprehension eagerly loads all 50,000,00050{,}000{,}000 images into RAM before training even begins, immediately crashing the system with an OutOfMemoryError.

  1. Functions, *args, **kwargs, Lambdas, and Decorators

Functions in Python are First-Class Objects—meaning you can pass a loss function or activation function as an argument into another function, return functions, and wrap them with decorators:

1. Positional, Default & Keyword Arguments

Functions accept required positional parameters followed by optional default parameters (e.g., def train(model, epochs=10, lr=1e-3):). Passing keyword arguments explicitly (train(net, lr=5e-4)) prevents ordering bugs.

2. *args (Tuple) and **kwargs (Dict)

*args packs any number of extra positional arguments into a Tuple, while **kwargs packs extra named keyword arguments into a Dictionary. Used heavily when wrapping PyTorch layers and HuggingFace model configs!

3. Lambda Functions (Anonymous Inline)

Single-expression inline functions defined with lambda x: expr. Most commonly used as custom sorting keys—such as sorting predicted classes by probability: sorted(preds, key=lambda item: item[1], reverse=True).

4. Decorators (@torch.no_grad)

A decorator @wrapper is a higher-order function that wraps another function to modify its behavior without changing its source code—such as @torch.no_grad() to disable gradient tracking during LLM inference!

⚡ Knowledge Check

What happens if you define a function with a mutable list as a default parameter: def addMetric(val, history=[]): history.append(val); return history and call addMetric(0.9) twice without passing history?

A) The second call returns [0.9, 0.9] because default lists are created only ONCE▼
✓ Correct!Python evaluates default arguments only once when def is executed, not on every call! All calls share the exact same list in memory. Always use history=None and initialize if history is None: history = [] inside the function.
B) It creates a brand-new empty list [] on every call▼
✕ Incorrect.Mutable default arguments persist across function calls and are one of the most notorious bugs in Python data pipelines.

  1. Visual Explanation: How Data Structures Power an AI Pipeline

Look at how all four Python data structures and generator functions work together inside a real-world NLP Tokenization and Mini-Batch Data Loader pipeline before handing off to the GPU:

  1. Python Implementation: Building an AI Tokenizer & Batch Generator

Here is a complete, runnable pure-Python implementation combining Sets, Dict Comprehensions, Type Hints, Tuples, **kwargs, and a Lazy Generator to tokenize text and stream mini-batches:

ai_data_pipeline.pyPython 3.11+ · Core Data Structures
# 1. Build a Vocabulary using a Set (Unique O(1)) and Dict Comprehensioncorpus = [    "neural networks learn representations",    "transformers use attention networks",    "attention is all you need"] # Extract unique words via Set Comprehension, then sortuniqueWords = sorted({word for sentence in corpus for word in sentence.split()}) # Map word -> integer ID (O(1) Hash Map) with <PAD>=0 and <UNK>=1wordToId = {"<PAD>": 0, "<UNK>": 1}wordToId.update({w: idx + 2 for idx, w in enumerate(uniqueWords)})idToWord = {idx: w for w, idx in wordToId.items()} # 2. Function with Type Hints, Default Args (Safe None Pattern), and **kwargsdef encodeSentence(text: str, vocab: dict, specialTokens=None, **kwargs) -> tuple:    if specialTokens is None:        specialTokens = []                      # Safe mutable default initialization!    maxLen = kwargs.get("maxLen", 6)    tokenIds = [vocab.get(w, 1) for w in text.split()][:maxLen]    padding = [0] * max(0, maxLen - len(tokenIds))    finalIds = tokenIds + padding    return (finalIds, len(tokenIds))            # Return immutable Tuple (ids, validLength) # 3. Lazy Mini-Batch Generator using yield (O(1) Memory)def batchLoader(sentences: list, vocab: dict, batchSize: int = 2):    for i in range(0, len(sentences), batchSize):        batchSlice = sentences[i : i + batchSize]        encodedBatch = [encodeSentence(s, vocab, maxLen=5) for s in batchSlice]        yield encodedBatch # 4. Stream Batches and Unpack Tuplesfor step, batch in enumerate(batchLoader(corpus, wordToId, batchSize=2)):    print(f"Step {step} -> Batch of {len(batch)} samples:", batch)

Pro Tip (Safe Dictionary Lookups with .get()): Notice how we wrote vocab.get(w, 1) instead of vocab[w]! If a user types an unseen out-of-vocabulary word at inference time, vocab[w] crashes with a KeyError, whereas vocab.get(w, 1) safely falls back to the <UNK> token ID 1.

Key Points

✓Python variables are references to objects in memory; mutable objects (Lists, Dicts, Sets) can be modified in-place, while immutable objects (Tuples, Strings, Ints, Floats) cannot.
✓Dictionaries and Sets use hash tables to achieve O(1)\mathcal{O}(1) constant-time lookups, making them thousands of times faster than Lists (O(N)\mathcal{O}(N)) for vocabulary mapping and membership checks.
✓Tuples are immutable and hashable, making them the standard choice for fixed tensor shapes (Batch, Seq, Dim) and multi-value function returns.
✓List and Dict Comprehensions provide concise, fast C-level loops for transforming data, while Generators (yield) stream items lazily one batch at a time in O(1)\mathcal{O}(1) RAM.
✓Functions in Python are first-class objects that support flexible argument packing via *args (tuple) and **kwargs (dictionary), inline lambda expressions, and @decorators.

Common Mistakes

✕ Using a mutable list or dict as a default function argument (def fn(x, cache=[]):).

Default arguments are evaluated only once when the function is defined. Mutating that default list leaks state across separate function calls! Always default to None and create a new list inside the function.

✕ Copying a list with b = a and assuming changes to b will not affect a.

Assignment only copies the memory reference, not the underlying list. Use a.copy() for 1D lists or copy.deepcopy(a) for nested structures.

✕ Creating a 2D matrix in Python using [[0] * cols] * rows.

Multiplying a nested list copies the reference to the same inner row NN times—so setting grid[0][0] = 5 changes column 00 in every single row! Always use [[0] * cols for r in range(rows)] or NumPy arrays.

✕ Using pure Python for-loops to multiply large numerical matrices.

Pure Python lists store boxed objects scattered across memory and check types on every iteration. Use Python lists/dicts for orchestration and I/O, and offload heavy math to vectorized NumPy or PyTorch tensors.

The Big Picture

Naive Scripting (Slow & Memory-Hungry)

Load Entire Dataset in List → O(N) Linear Searches → Copy Bugs → Out-of-Memory Crash

Production AI Python (Clean & Scalable)

O(1) Dict/Set Vocabularies → Lazy yield Generators → Modular Typed Functions → GPU Tensors

The important conceptual shift is understanding how Python divides labor in AI: built-in data structures (Dicts, Sets, Tuples, Lists) and lazy Generators handle all the messy text parsing, JSON configuration, and data streaming on the CPU, feeding clean, structured batches into vectorized C++/CUDA libraries like NumPy and PyTorch.

Remember: Write Pythonic orchestration code—use Sets and Dicts for instant O(1)\mathcal{O}(1) lookups, Tuples for fixed shapes, and Generators for streaming data—before passing your numbers into NumPy and PyTorch for heavy matrix math.