PyTorch Tensors and Automatic Differentiation

From beginner GPU tensors to medium tensor reshaping and advanced Autograd computational graphs that power modern Deep Learning.

27 minIntermediateCode Examples

The Core Thesis: NumPy is great for fast math on a CPU, but it lacks two superpowers required for Deep Learning: it cannot run on GPUs, and it cannot automatically calculate derivatives. PyTorch solves both problems at once. A PyTorch Tensor is a multi-dimensional array that can live directly on a GPU and record every mathematical operation performed on it—allowing PyTorch's Autograd engine to compute millions of partial derivatives automatically in a single backward pass.

  1. Beginner Foundations: What Makes a PyTorch Tensor Special?

On the surface, a PyTorch Tensor (torch.Tensor) looks almost identical to a NumPy array. Every tensor has three fundamental properties that you should always check when debugging a neural network:

1. Shape (x.shape)
torch.Size([64, 128])

Tells you the size along each dimension—such as 6464 batch samples by 128128 features.

2. Data Type (x.dtype)
float32, bfloat16, int64

Model weights use torch.float32 (or bfloat16 for LLMs), while token IDs and class labels require torch.int64 (torch.long).

3. Device (x.device)
"cpu", "cuda:0", "mps"

Tells you where the tensor lives physically: system RAM (cpu), an NVIDIA GPU (cuda), or Apple Silicon (mps).

Zero-Copy NumPy Bridge: On the CPU, torch.from_numpy(arr) and tensor.numpy() share the exact same underlying memory address without copying data! That makes converting cleaned Pandas/NumPy data into PyTorch instantaneous.

  1. Intermediate: Reshaping, Permuting, and Broadcasting Tensors

In Deep Learning, 80%80\% of writing model architectures (like CNNs and Multi-Head Attention) is reshaping tensors so their dimensions align for matrix multiplication (@). Here are the essential shape tools:

OperationPyTorch SyntaxWhat It DoesExample Shape Change
View / Reshapex.view(B, -1) or x.reshape(B, -1)Changes tensor dimensions without changing total element count (-1 infers size automatically).(32, 3, 28, 28) → (32, 2352)
Unsqueezex.unsqueeze(0)Inserts a new dimension of size 11 (used to add a batch dimension to a single sample).(3, 224, 224) → (1, 3, 224, 224)
Squeezex.squeeze(-1)Removes a dimension of size 11 (e.g., turning column predictions into a 1D vector).(64, 1) → (64,)
Transpose / Permutex.transpose(1, 2) / x.permute(0, 2, 1)Swaps axes (used to split and reorder Multi-Head Attention heads).(B, Seq, Heads, Dim) → (B, Heads, Seq, Dim)
Batched MatMulQ @ K.transpose(-2, -1)Multiplies matrices across the last two dimensions for every batch and head in parallel.(B, H, T, d) @ (B, H, d, T) → (B, H, T, T)
⚡ Knowledge Check

Suppose your neural network weights W are on the GPU (device="cuda:0"), and you load a new user input tensor x on the CPU (device="cpu") with shape (768,). What happens if you run x @ W directly?

A) RuntimeError: Tensors must be on the same device!▼
✓ Correct!PyTorch never moves data between CPU RAM and GPU VRAM behind your back because PCIe transfers are slow. You must explicitly move x = x.to("cuda:0") so both tensors live on the exact same device.
B) PyTorch automatically copies x to the GPU and returns the result▼
✕ Incorrect.Mixing CPU and CUDA tensors is one of the top 3 runtime errors in PyTorch: RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!

  1. Intermediate: How Automatic Differentiation (Autograd) Works

In our Calculus lessons, we calculated partial derivatives ∂L∂w\dfrac{\partial L}{\partial w} by hand using the Chain Rule. When a model has 100100 million parameters, doing derivatives by hand is impossible. Instead, PyTorch uses Reverse-Mode Automatic Differentiation (Autograd):

1. Turn On Tracking (requires_grad=True)

When you create a learnable weight tensor with requires_grad=True, PyTorch starts recording every single operation applied to that tensor in a Dynamic Computational Graph (DAG).

w = torch.tensor([3.0], requires_grad=True)

2. Run Backward Pass (loss.backward())

When you call loss.backward() on the final scalar Loss, PyTorch walks backward through the recorded graph using the Chain Rule and stores the exact derivative ∂L∂w\dfrac{\partial L}{\partial w} inside w.grad!

loss.backward() # Populates w.grad!

Under the Hood: Leaf Tensors vs. grad_fn Pointers

Suppose x=2x = 2, weight w=3w = 3 (requires_grad=True), bias b=1b = 1 (requires_grad=True), and we compute y=wx+by = wx + b followed by L=(y−10)2L = (y - 10)^2:

Leaf Tensors (w and b)
Created directly by the user. They are the leaves of the graph where PyTorch accumulates w.grad (−12.0-12.0) and b.grad (−6.0-6.0).
Intermediate Tensors (y and L)
Created by math ops. Each stores a .grad_fn pointer (like <AddBackward0> or <PowBackward0>) telling Autograd which derivative formula to run!

  1. Intermediate: Why Gradients Accumulate (zero_grad) & Disabling Autograd

Every time you call loss.backward(), PyTorch ADDS (accumulates) the new gradient into w.grad instead of overwriting it:

w.gradnew=w.gradold+∂L∂w\text{w.grad}_{\text{new}} = \text{w.grad}_{\text{old}} + \dfrac{\partial L}{\partial w}

Why did PyTorch design it this way? Because of the Multivariable Chain Rule (when a weight is used in two places, its gradients must sum) AND so engineers with small GPUs can run Gradient Accumulation (running 44 small micro-batches of size 88 before updating weights to simulate a large batch of 3232)!

optimizer.zero_grad()

Resets all .grad buffers back to 00 before the next training step so old gradients do not corrupt your new batch.

with torch.no_grad():

Temporarily turns off graph recording during validation, testing, and manual weight updates—cutting GPU memory usage in half!

tensor.detach()

Returns a new tensor sharing the same data, but detached from the computational graph (needed before converting to NumPy/Matplotlib).

⚡ Knowledge Check

Suppose w = torch.tensor(3.0, requires_grad=True) and L=w2L = w^2 (so dLdw=2w=6\dfrac{dL}{dw} = 2w = 6). If you accidentally call L = w ** 2; L.backward() THREE times in a row without clearing gradients, what value will be stored inside w.grad?

A) 18.0 (because 6.0 + 6.0 + 6.0 accumulates!)▼
✓ Correct!Every call to .backward() adds +6.0+6.0 onto the existing w.grad buffer (6+6+6=18.06 + 6 + 6 = 18.0). Always clear gradients between steps using w.grad.zero_() or optimizer.zero_grad().
B) 6.0 (because backward overwrites w.grad)▼
✕ Incorrect.PyTorch never overwrites .grad automatically during .backward(); it always accumulates via addition.

  1. Advanced: Memory Strides (.contiguous), In-Place Traps, and Mixed Precision

When building production Transformers and training loops, three advanced PyTorch mechanics are essential to master:

1. Memory Strides: .view() vs .contiguous()

When you transpose a tensor (x.transpose(1, 2)), PyTorch does not move numbers in memory—it only swaps the metadata strides! Calling .view() on a transposed tensor crashes because the memory is no longer contiguous.

Fix: x.transpose(1, 2).contiguous().view(...) # Or use .reshape()

2. The In-Place Operation Trap (add_, mul_)

In PyTorch, any method ending with an underscore (like x.add_(1) or x += 1) modifies the tensor in-place, overwriting its original memory. Doing this on an active autograd tensor destroys the forward values needed for the Chain Rule!

RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation

3. Advanced LLM Speedup (Mixed Precision bfloat16 & torch.inference_mode()): Modern GPUs (Ampere/Hopper Tensor Cores) multiply 16-bit Brain Floating Point (torch.bfloat16) matrices up to 4×4\times faster than 32-bit floats while using half the VRAM. For production inference, prefer with torch.inference_mode(): over torch.no_grad()—it completely disables view tracking and version counters for maximum speed!

⚡ Knowledge Check

When writing a custom Gradient Descent weight update on a leaf tensor w (requires_grad=True), why must you wrap the update inside with torch.no_grad(): w -= lr * w.grad instead of writing w = w - lr * w.grad outside no_grad?

A) So the weight update itself is not recorded in the graph, keeping w as a leaf tensor▼
✓ Correct!Writing w = w - lr * w.grad with autograd enabled creates a brand-new non-leaf tensor connected to the old graph, leaking memory across steps and breaking future w.grad access!
B) Because PyTorch cannot subtract two tensors without no_grad▼
✕ Incorrect.PyTorch subtracts tensors normally during the forward pass; we use torch.no_grad() during optimizer updates specifically so the optimizer step is not tracked as part of the neural network's forward math.

  1. Visual Explanation: The Dynamic Autograd Computational Graph

Look at how PyTorch builds the Dynamic Computational Graph on the fly during the Forward Pass, attaches a grad_fn backward pointer to every intermediate tensor, and walks backward on loss.backward() to populate the leaf gradients W.grad and b.grad[cite: 12]:

  1. Python Implementation: Complete Autograd Training Loop From Scratch

Here is a complete, runnable PyTorch script showing Hardware-Agnostic Device Selection, Leaf Parameter Tensors, Automatic Differentiation (loss.backward()), and a clean torch.no_grad() Optimization Loop that trains a linear model y=3x1−2x2+5y = 3x_1 - 2x_2 + 5 from scratch:

pytorch_autograd_loop.pyPython 3.11+ · PyTorch 2.x
import torch # 1. Hardware-Agnostic Device Setup (CUDA GPU -> Apple MPS -> CPU)device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")torch.manual_seed(42) # 2. Create Synthetic Dataset on Device: True Formula y = 3x1 - 2x2 + 5X = torch.randn(64, 2, device=device)trueW = torch.tensor([[3.0], [-2.0]], device=device)trueB = torch.tensor([5.0], device=device)Y = X @ trueW + trueB # 3. Initialize Learnable Leaf Tensors with requires_grad=TrueW = torch.randn(2, 1, device=device, requires_grad=True)b = torch.zeros(1, device=device, requires_grad=True)lr = 0.1 # 4. Training Loop Powered by PyTorch Autogradfor epoch in range(1, 61):    # Forward Pass (Builds dynamic computational graph automatically)    yPred = X @ W + b    loss = torch.mean((yPred - Y) ** 2)     # Backward Pass (Computes dL/dW and dL/db via Reverse-Mode Autodiff)    loss.backward()     # Update Weights inside torch.no_grad() and Reset Gradients to Zero!    with torch.no_grad():        W -= lr * W.grad        b -= lr * b.grad        W.grad.zero_()        b.grad.zero_() print("Learned W:", W.detach().cpu().numpy().round(3).flatten())  # [3.0, -2.0]print("Learned b:", round(float(b.item()), 3))                    # 5.0

Pro Tip (Why We Use .item() for Logging Loss): When printing or storing training loss in a Python list (loss_history.append(loss.item())), always call .item()! Appending the raw loss tensor directly into a list keeps the entire computational graph attached in GPU memory for every step until your GPU crashes with CUDA out of memory.

Key Points

✓A PyTorch Tensor is a multi-dimensional array defined by its shape, dtype (float32/bfloat16 for weights, int64 for indices), and device (cpu, cuda, mps).
✓Use .view() or .reshape() to change tensor shapes, .unsqueeze() / .squeeze() to add or remove size-11 dimensions, and .transpose() / .permute() to reorder axes.
✓Setting requires_grad=True on leaf parameter tensors tells PyTorch's Autograd engine to record a dynamic computational graph of all forward operations.
✓Calling loss.backward() traverses the computational graph in reverse using the Chain Rule and accumulates exact partial derivatives inside each leaf tensor's .grad attribute.
✓Because PyTorch accumulates gradients by addition, you must zero out gradients (zero_() or optimizer.zero_grad()) on every training step.
✓Wrap validation loops and manual weight updates inside with torch.no_grad(): (or torch.inference_mode()) to disable graph construction and save GPU memory.

Common Mistakes

✕ Forgetting to zero out gradients between batches (missing zero_grad).

If you do not clear .grad after updating weights, the next batch's gradients are added onto the old batch's gradients, causing weights to overshoot wildly.

✕ Appending the raw loss tensor (history.append(loss)) instead of loss.item().

Storing a tensor that has a grad_fn inside a Python list prevents Python's garbage collector from freeing the entire forward activation graph, quickly crashing your GPU with an Out-Of-Memory error.

✕ Mixing CPU and GPU tensors in the same operation.

Both the model parameters and the input batch must live on the exact same device (e.g., batch = batch.to(device)) before running the forward pass.

✕ Calling .view() directly after .transpose() without .contiguous() or .reshape().

Transposing a tensor changes its stride metadata without rearranging physical memory. Call .contiguous().view(...) or simply .reshape(...) to avoid stride errors.

The Big Picture

Manual NumPy Code (CPU Only & Hand-Written Calculus)

CPU Arrays → Derive Every Layer's Gradient Formula by Hand → Slow & Error-Prone

PyTorch Tensors + Autograd (GPU Acceleration & Automatic Calculus)

Write Forward Math on GPU → Autograd Records Graph → loss.backward() Computes All Gradients Automatically

The important conceptual shift is realizing that with PyTorch, you only ever have to write the Forward Pass. As long as you compose your model out of differentiable PyTorch tensor operations, the Autograd engine automatically builds the backward graph and computes every partial derivative for you on the GPU.

Remember: Every modern deep learning model—from a 2-layer classifier to GPT-4—boils down to the exact 4-step PyTorch rhythm you just learned: Forward Pass → loss.backward() → Update Weights → zero_grad().