Matrix Multiplication and Linear Transformations

How matrices act as geometric functions that rotate, scale, project, and warp vector spaces—forming the mathematical backbone of every neural network layer.

24 minBeginnerCode Examples

The Core Thesis: A matrix is not just a passive spreadsheet of numbers—it is an active machine that transforms space. Every time you multiply a vector by a matrix (y=Ax\mathbf{y} = \mathbf{A}\mathbf{x}), you are stretching, rotating, or projecting the entire coordinate grid to extract new features from raw data.

  1. What is a Linear Transformation?

In algebra, a function f(x)f(x) takes a scalar number as input and outputs another number. In Linear Algebra, a Transformation T(x)T(\mathbf{x}) is a function that takes an entire input vector x\mathbf{x} and maps it to a new output vector y\mathbf{y}:

T(x)=Ax=yT(\mathbf{x}) = \mathbf{A}\mathbf{x} = \mathbf{y}

Why is it called Linear? Visually, a transformation is strictly linear if and only if it obeys two geometric rules:

Rule 1: Grid Lines Remain Parallel

All straight lines in space stay straight after the transformation—they never curve or bend. Furthermore, parallel grid lines remain parallel and evenly spaced.

T(u+v)=T(u)+T(v)T(\mathbf{u} + \mathbf{v}) = T(\mathbf{u}) + T(\mathbf{v})
Rule 2: The Origin Remains Fixed

The center of the coordinate space (0,0)(0, 0) never shifts. If you pass a zero vector into a matrix multiplication, you always get the zero vector back.

T(cu)=cT(u)andA0=0T(c\mathbf{u}) = cT(\mathbf{u}) \quad \text{and} \quad \mathbf{A}\mathbf{0} = \mathbf{0}

AI Connection (Why Neural Networks Need Bias b\mathbf{b}): Because pure matrix multiplication Wx\mathbf{W}\mathbf{x} locks the origin at (0,0)(0,0), a neural network layer adds a bias vector b\mathbf{b} to create an Affine Transformation (y=Wx+b\mathbf{y} = \mathbf{W}\mathbf{x} + \mathbf{b}), allowing the decision boundary to shift away from the origin.

  1. The Secret of Basis Vectors: Reading Any Matrix Instantly

How can you look at a 2×22 \times 2 matrix like [3102]\begin{bmatrix} 3 & 1 \\ 0 & 2 \end{bmatrix} and immediately see what it does to space? The secret lies in the Standard Basis Vectors:

Horizontal Unit Vector (i^\hat{i})
Points 1 unit right along the X-axis
i^=[10]\hat{i} = \begin{bmatrix} 1 \\ 0 \end{bmatrix}
Vertical Unit Vector (j^\hat{j})
Points 1 unit up along the Y-axis
j^=[01]\hat{j} = \begin{bmatrix} 0 \\ 1 \end{bmatrix}

Every vector x=[xy]\mathbf{x} = \begin{bmatrix} x \\ y \end{bmatrix} is just a sum of scaled basis vectors: xi^+yj^x\hat{i} + y\hat{j}. When you multiply a matrix A=[abcd]\mathbf{A} = \begin{bmatrix} a & b \\ c & d \end{bmatrix} by x\mathbf{x}, look at what happens to the columns:

[abcd][xy]=x[ac]+y[bd]\begin{bmatrix} a & b \\ c & d \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} = x \begin{bmatrix} a \\ c \end{bmatrix} + y \begin{bmatrix} b \\ d \end{bmatrix}

Column 1 [ac]\begin{bmatrix} a \\ c \end{bmatrix} is literally where the X-axis basis vector i^=[1,0]T\hat{i} = [1, 0]^T lands!

Column 2 [bd]\begin{bmatrix} b \\ d \end{bmatrix} is literally where the Y-axis basis vector j^=[0,1]T\hat{j} = [0, 1]^T lands!

⚡ Knowledge Check

A 2D linear transformation moves the horizontal basis vector i^=[1,0]T\hat{i} = [1, 0]^T to [2,−1]T[2, -1]^T, and moves the vertical basis vector j^=[0,1]T\hat{j} = [0, 1]^T to [0,4]T[0, 4]^T. Without doing any algebra, what is the 2×22 \times 2 transformation matrix A\mathbf{A}?

A) First col [2, -1], Second col [0, 4]▼
✓ Correct!The columns of any transformation matrix are simply the landing coordinates of the basis vectors i^\hat{i} and j^\hat{j}, giving A=[20−14]\mathbf{A} = \begin{bmatrix} 2 & 0 \\ -1 & 4 \end{bmatrix}.
B) First row [2, -1], Second row [0, 4]▼
✕ Incorrect.Transformed basis vectors always populate the columns of the matrix, not the rows.

  1. The Visual Catalog of Geometric Matrix Transformations

By changing the entries of a matrix A\mathbf{A}, we can perform specific geometric operations on data points (such as image augmentation in Computer Vision or feature projection in Transformers):

1. Scaling (Stretch / Shrink)Diagonal Matrix

Stretches X by factor sxs_x and Y by factor sys_y. Off-diagonal zeros mean axes do not mix.

A=[sx00sy]=[2000.5]\mathbf{A} = \begin{bmatrix} s_x & 0 \\ 0 & s_y \end{bmatrix} = \begin{bmatrix} 2 & 0 \\ 0 & 0.5 \end{bmatrix}
2. Shear (Slanting Space)Upper Triangular

Keeps i^\hat{i} fixed at [1,0]T[1,0]^T, but pushes j^\hat{j} sideways to [k,1]T[k, 1]^T, slanting the grid like italics.

A=[1k01]=[11.501]\mathbf{A} = \begin{bmatrix} 1 & k \\ 0 & 1 \end{bmatrix} = \begin{bmatrix} 1 & 1.5 \\ 0 & 1 \end{bmatrix}
3. Counter-Clockwise RotationOrthogonal Matrix

Rotates every vector by angle θ\theta around the origin without changing lengths (powers RoPE in LLMs!).

R(θ)=[cos⁡θ−sin⁡θsin⁡θcos⁡θ]\mathbf{R}(\theta) = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix}
4. Reflection & ProjectionRank-Deficient

Flips space across an axis or squashes 2D space onto a 1D line by zeroing out a dimension.

Px-axis=[1000]\mathbf{P}_{\text{x-axis}} = \begin{bmatrix} 1 & 0 \\ 0 & 0 \end{bmatrix}

  1. Two Ways to Compute Matrix Multiplication

When multiplying an (m×k)(m \times k) matrix A\mathbf{A} by a (k×n)(k \times n) matrix B\mathbf{B} to produce C\mathbf{C} of shape (m×n)(m \times n), mathematicians and AI engineers use two complementary mental models:

1. The Row-by-Column Dot Product View

Each individual cell ci,jc_{i,j} is computed by taking the dot product of Row ii from A\mathbf{A} and Column jj from B\mathbf{B}.

ci,j=∑p=1kai,pbp,jc_{i,j} = \sum_{p=1}^{k} a_{i,p} b_{p,j}

Best for: Computing single entries by hand and checking feature-to-neuron similarity scores.

2. The Column Combination View

Each column of C\mathbf{C} is a linear combination of the columns of A\mathbf{A}, weighted by the numbers in that column of B\mathbf{B}.

Abj=b1,ja:,1+b2,ja:,2+⋯+bk,ja:,k\mathbf{A}\mathbf{b}_j = b_{1,j}\mathbf{a}_{:,1} + b_{2,j}\mathbf{a}_{:,2} + \dots + b_{k,j}\mathbf{a}_{:,k}

Best for: Understanding how Attention blends Value vectors and how Low-Rank Adaptation (LoRA) works.

Step-by-Step Worked Example (Both Views Yield the Same Answer):

Multiply A=[1324]\mathbf{A} = \begin{bmatrix} 1 & 3 \\ 2 & 4 \end{bmatrix} by vector x=[52]\mathbf{x} = \begin{bmatrix} 5 \\ 2 \end{bmatrix}:
Row Dot-Product Method
[(1⋅5+3⋅2)(2⋅5+4⋅2)]=[5+610+8]=[1118]\begin{bmatrix} (1\cdot 5 + 3\cdot 2) \\ (2\cdot 5 + 4\cdot 2) \end{bmatrix} = \begin{bmatrix} 5 + 6 \\ 10 + 8 \end{bmatrix} = \begin{bmatrix} 11 \\ 18 \end{bmatrix}
Column Combination Method
5[12]+2[34]=[510]+[68]=[1118]5 \begin{bmatrix} 1 \\ 2 \end{bmatrix} + 2 \begin{bmatrix} 3 \\ 4 \end{bmatrix} = \begin{bmatrix} 5 \\ 10 \end{bmatrix} + \begin{bmatrix} 6 \\ 8 \end{bmatrix} = \begin{bmatrix} 11 \\ 18 \end{bmatrix}

  1. Matrix Multiplication as Function Composition

Why do we multiply two matrices B\mathbf{B} and A\mathbf{A} together? Because if matrix A\mathbf{A} applies one transformation (say, a Rotation) and matrix B\mathbf{B} applies a second transformation (say, a Shear), then the product matrix C=BA\mathbf{C} = \mathbf{B}\mathbf{A} is a single combined transformation that performs both in sequence!

Right-to-Left Execution Order (Just Like Nested Functions)

y=B(Ax)=(BA)x\mathbf{y} = \mathbf{B}(\mathbf{A}\mathbf{x}) = (\mathbf{B}\mathbf{A})\mathbf{x}

Because x\mathbf{x} is on the right, Matrix A\mathbf{A} acts on x\mathbf{x} FIRST, and Matrix B\mathbf{B} acts on the result SECOND.

This immediately explains why matrix multiplication is NOT commutative (BA≠AB\mathbf{B}\mathbf{A} \neq \mathbf{A}\mathbf{B}): rotating an image by 90∘90^\circ and then shearing it produces a completely different geometric result than shearing it first and rotating it second. However, it IS associative: (CB)A=C(BA)(\mathbf{C}\mathbf{B})\mathbf{A} = \mathbf{C}(\mathbf{B}\mathbf{A}).

⚡ Knowledge Check

Suppose you stack 3 linear neural network layers W1\mathbf{W}_1, W2\mathbf{W}_2, and W3\mathbf{W}_3 directly after one another WITHOUT any non-linear activation functions (like ReLU) in between: y=W3(W2(W1x))\mathbf{y} = \mathbf{W}_3(\mathbf{W}_2(\mathbf{W}_1\mathbf{x})). What happens mathematically?

A) All 3 layers collapse into a single matrix W_total▼
✓ Correct!By associativity, W3(W2(W1x))=(W3W2W1)x=Wsinglex\mathbf{W}_3(\mathbf{W}_2(\mathbf{W}_1\mathbf{x})) = (\mathbf{W}_3\mathbf{W}_2\mathbf{W}_1)\mathbf{x} = \mathbf{W}_{\text{single}}\mathbf{x}. Without non-linear activations (ReLU/GELU) between matrices, a 100-layer deep network has the exact same expressive power as a 1-layer linear model!
B) It learns complex non-linear curves▼
✕ Incorrect.Multiplying linear matrices together always produces another strictly linear matrix. You must insert non-linear activation functions between matrix multiplications to bend space and learn curves.

  1. Dimensionality Changes: Expanding and Compressing Space

Matrices do not have to be square! A non-square matrix W\mathbf{W} of shape (m×n)(m \times n) maps vectors from nn-dimensional input space (Rn\mathbb{R}^n) into mm-dimensional output space (Rm\mathbb{R}^m):

Tall Matrix (m > n): Up-Projection

Maps low-dimensional inputs into a higher-dimensional feature space. In Transformer Feed-Forward networks, hidden vectors of size 768768 are multiplied by a (3072×768)(3072 \times 768) matrix to expand into 30723072 dimensions.

Shape: (3072, 768) · (768, 1) = (3072, 1)

Wide Matrix (m < n): Down-Projection

Compresses high-dimensional vectors down into a lower-dimensional bottleneck (used in Autoencoders, PCA, and Multi-Head Attention query projections).

Shape: (64, 768) · (768, 1) = (64, 1)

Determinants, Rank, and Invertibility

When a square matrix A\mathbf{A} transforms space, how much does the area (or volume) of a region scale? That scaling factor is called the Determinant, written det⁡(A)\det(\mathbf{A}):

det⁡([abcd])=ad−bc\det\left(\begin{bmatrix} a & b \\ c & d \end{bmatrix}\right) = ad - bc
det⁡(A)≠0\det(\mathbf{A}) \neq 0 (Full Rank)

Space is stretched or rotated without being squashed flat. The transformation is reversible via the Inverse Matrix A−1\mathbf{A}^{-1} (A−1A=I\mathbf{A}^{-1}\mathbf{A} = \mathbf{I}).

\det(\mathbf{A}) < 0 (Flipped Space)

The orientation of space is inverted (like flipping a sheet of paper over), while areas scale by ∣det⁡(A)∣|\det(\mathbf{A})|.

det⁡(A)=0\det(\mathbf{A}) = 0 (Singular)

Columns are linearly dependent! Space is squashed into a lower dimension (e.g., 2D plane collapsed onto a 1D line). No inverse exists; information is permanently lost.

⚡ Knowledge Check

Consider the matrix A=[2412]\mathbf{A} = \begin{bmatrix} 2 & 4 \\ 1 & 2 \end{bmatrix}. Notice that Column 2 is simply 2×2 \times Column 1. What is det⁡(A)\det(\mathbf{A}), and can you recover an input vector x\mathbf{x} after multiplying y=Ax\mathbf{y} = \mathbf{A}\mathbf{x}?

A) det(A) = 0; Impossible to invert (Information Lost)▼
✓ Correct!det⁡(A)=(2)(2)−(4)(1)=0\det(\mathbf{A}) = (2)(2) - (4)(1) = 0. Because both basis vectors land on the exact same line, the entire 2D plane is crushed onto a 1D line. Infinitely many inputs x\mathbf{x} map to the same output y\mathbf{y}, so A−1\mathbf{A}^{-1} does not exist.
B) det(A) = 8; Fully invertible via A^-1▼
✕ Incorrect.Remember the formula is ad−bc=(2×2)−(4×1)=0ad - bc = (2 \times 2) - (4 \times 1) = 0, not ad+bcad + bc. Whenever one column is a multiple of another, the determinant is zero and the matrix is non-invertible (singular).

  1. Visual Explanation: Neural Layers as Space Warping

What is a Deep Neural Network actually doing geometrically? Each layer applies a Linear Transformation (W\mathbf{W}) to rotate and stretch the data space, shifts it with a Bias (b\mathbf{b}), and folds it with a Non-Linear Activation (σ\sigma) until tangled classes become cleanly separable by a flat hyperplane:

Modern AI MechanismLinear Transformation FormulaGeometric Purpose
Dense Layer (nn.Linear)y = x @ W.T + bProjects dind_{\text{in}} features into doutd_{\text{out}} learned representations.
Transformer Q, K, V ProjectionsQ = X @ W_Q, K = X @ W_K, V = X @ W_VMaps a single token embedding into 3 distinct semantic subspaces (Query, Key, Value).
Rotary Position Embedding (RoPE)q_rot = R(m * theta) @ qRotates token vectors in 2D planes by an angle proportional to sequence position mm.
LoRA (Low-Rank Adaptation)W_new = W_frozen + (B @ A)Compresses weight updates through a tiny rank-rr bottleneck (d→r→dd \to r \to d) to fine-tune LLMs cheaply.

  1. Python Implementation: Geometric Transformations & Projections

Below is a complete NumPy implementation showing how to rotate vectors in 2D space, check matrix invertibility via the determinant, and project high-dimensional embeddings into a lower subspace:

linear_transformations.pyPython 3 · NumPy
import numpy as np # 1. Geometric 2D Rotation Matrix (Rotate 90 degrees counter-clockwise)theta = np.radians(90)R = np.array([    [np.cos(theta), -np.sin(theta)],    [np.sin(theta),  np.cos(theta)]]) v = np.array([3.0, 0.0])                    # Vector pointing along X-axisv_rotated = R @ v                           # Apply linear transformationprint("Rotated Vector:", np.round(v_rotated, 2))  # Output: [0. 3.] (Now on Y-axis!) # 2. Determinant & Matrix Inverse CheckA = np.array([[2.0, 1.0], [1.0, 3.0]])det_A = np.linalg.det(A)print("Determinant of A:", round(det_A, 2))       # Output: 5.0 (Invertible!) A_inv = np.linalg.inv(A)                    # Undo transformation: A_inv @ (A @ v) == vrecovered_v = A_inv @ (A @ v)print("Recovered Original Vector:", recovered_v)  # Output: [3. 0.] # 3. Dimensionality Projection (Transformer Query Projection)# Batch of 4 tokens, each a 512-D embedding -> Project down to 64-D headX_tokens = np.random.randn(4, 512)W_Query = np.random.randn(512, 64) * 0.02Q_projected = X_tokens @ W_Query            # Shape: (4, 512) @ (512, 64) -> (4, 64)print("Projected Query Shape:", Q_projected.shape)

Pro Tip (Convention Difference): In math textbooks, a single vector x\mathbf{x} is a column of shape (din,1)(d_{\text{in}}, 1), so we write y=Wx\mathbf{y} = \mathbf{W}\mathbf{x}. In PyTorch and NumPy, a batch of data X\mathbf{X} stores samples as rows of shape (Batch,din)(\text{Batch}, d_{\text{in}}), so we multiply on the right: Y = X @ W.T!

Key Points

✓Every matrix A\mathbf{A} represents a linear transformation that maps input vectors to output vectors while keeping grid lines straight, parallel, and evenly spaced.
✓The columns of any transformation matrix tell you the exact coordinates where the basis unit vectors (i^,j^,…\hat{i}, \hat{j}, \dots) land after transformation.
✓Multiplying two matrices (C=BA\mathbf{C} = \mathbf{B}\mathbf{A}) composes their geometric transformations from right to left: A\mathbf{A} transforms space first, followed by B\mathbf{B}.
✓Non-square matrices (m×n)(m \times n) change the dimensionality of data, projecting vectors between Rn\mathbb{R}^n and Rm\mathbb{R}^m (powering embeddings, bottlenecks, and Attention heads).
✓The determinant det⁡(A)\det(\mathbf{A}) measures how much a square matrix scales area or volume; if det⁡(A)=0\det(\mathbf{A}) = 0, space collapses into a lower dimension and information is irreversibly lost.
✓Deep neural networks alternate between linear matrix transformations (Wx+b\mathbf{W}\mathbf{x} + \mathbf{b}) and non-linear activations (ReLU/GELU) so they do not collapse into a single linear matrix.

Common Mistakes

✕ Reading matrix transformations from left to right instead of right to left.

In the expression CBAx\mathbf{C}\mathbf{B}\mathbf{A}\mathbf{x}, matrix A\mathbf{A} multiplies x\mathbf{x} first, then B\mathbf{B}, and finally C\mathbf{C}. Reversing this order produces an entirely different geometric transformation.

✕ Stacking linear layers without non-linear activation functions.

Because matrix multiplication is associative, stacking W2(W1x)\mathbf{W}_2(\mathbf{W}_1\mathbf{x}) without an activation function in between simply collapses into a single matrix Wcombinedx\mathbf{W}_{\text{combined}}\mathbf{x}, preventing the network from learning non-linear patterns.

✕ Calling np.linalg.inv(A) directly in production ML code.

Explicitly inverting large matrices is slow (O(n3)\mathcal{O}(n^3)) and numerically unstable when det⁡(A)\det(\mathbf{A}) is close to zero. Always use linear system solvers like np.linalg.solve(A, b) or gradient-based optimization instead.

✕ Confusing mathematical column-vector notation (W @ x) with batched row-vector code (X @ W.T).

Textbooks treat a single sample as a column vector (d,1)(d, 1), whereas PyTorch and NumPy stack batches as rows (B,d)(B, d). Always check whether your weight matrix needs to be transposed when moving from paper equations to code.

The Big Picture

Static Spreadsheet View (Arithmetic)

Grid of Numbers → Tedious Row-by-Column Sums → Another Grid of Numbers

Geometric Transformation View (Deep Learning)

Input Vector Space → Rotate, Scale & Project via Matrix W → Separable Feature Space

The important conceptual shift is realizing that training a neural network—whether a simple classifier or a trillion-parameter Transformer—is literally the process of learning the exact matrix entries needed to warp messy, tangled input vectors into a clean geometric space where answers are easy to separate.

Remember: Whenever you see a weight matrix W\mathbf{W} inside a neural network diagram, do not just see a table of parameters. See a geometric lens that rotates, compresses, or expands high-dimensional space to spotlight the features that matter.