TT Lab
Get started
Learn Learning paths Courses

3D Math and a Software Rasterizer

Why a 4x4 Matrix for Three Dimensions

Continue in TT Lab

In one line

Translation is addition, not multiplication, so it does not fit into a 3x3 matrix. If you add one more coordinate to make (x, y, z, 1), you can express even translation with a single multiplication, and that is why transformation matrices in graphics are 4x4.

Why this was needed

Putting one object on the screen takes a series of steps. You shrink it, rotate it, move it into place, move it again relative to the camera, and project it. If you chain these together as function calls, five different computations are attached to every vertex. With a hundred thousand vertices, that is five hundred thousand computations.

Things change when all of these computations can be expressed as a single matrix multiplication. Matrix multiplication is associative, so you can multiply the five together in advance into one, and per vertex you only need to do one 4x4 multiplication. This is why the GPU receives just one MVP matrix as a uniform. The problem is that translation is not a multiplication. A 3x3 matrix can do rotation and scaling but cannot express a translation.

How it works

The solution is to add one dimension. If you write the point (x, y, z) as (x, y, z, 1) and multiply by a 4x4 matrix, the last column is multiplied by that 1 and added to the result. If you put the translation amounts in that column, translation enters the multiplication.

| 1 0 0 tx |   | x |   | x + tx |
| 0 1 0 ty | * | y | = | y + ty |
| 0 0 1 tz |   | z |   | z + tz |
| 0 0 0  1 |   | 1 |   |   1    |

This representation is called homogeneous coordinates. The fourth component w is not just filler; it has a meaning. If w is 1 it is a point, and if it is 0 it is a direction. A translation must not be applied to a direction (north is still north wherever you move it), and if you set w=0 the translation column is multiplied by 0 and ignored automatically. That is why normal vectors and light directions are handled as (x, y, z, 0).

The order of multiplication changes the result. Matrix multiplication is not commutative. T * R means "rotate first, then translate", and R * T means "translate first, then rotate". In the convention where a column vector is multiplied on the right, the matrix on the right is applied first. If you apply a 90-degree rotation about the z axis and a translation of 2 in the x direction to the point (1,0,0), the former gives (2,1,0) and the latter gives (0,3,0). This is where it splits whether an object spins in place or orbits.

Rotation matrices have one more thing to get right: handedness. In a right-handed coordinate system, a rotation about the z axis turns the x axis toward the y axis.

rotate_z(t) = | cos t  -sin t  0  0 |
              | sin t   cos t  0  0 |
              |   0       0    1  0 |
              |   0       0    0  1 |

If you put in one sign wrong, the object rotates the opposite way. No error message appears.

There is also a difference in convention between row-major and column-major. This lab keeps a matrix as "four lists of length 4" and multiplies the point on the right. In contrast, DirectX-family documents and code often put the point on the left and multiply as a row vector, and then the matrix for the same transformation looks transposed. Between the two conventions, the multiplication order is reversed as well. When you port someone else's code and the result looks wrong, this is the first thing to suspect. The order in which a matrix is stored in memory is yet another axis, so you could count the conventions as four instead of two.

Matrices are not the only way to represent rotation. A quaternion holds a rotation in four values, and blending smoothly between two rotations (interpolation) works much better than with matrices. This is because blending two matrices component by component produces something that is not a rotation matrix. However, when handing data to the GPU you eventually have to convert to a matrix, so a common setup is to keep the state as a quaternion and build the matrix when drawing.

What it looks like in the field

If you use the model matrix as is to transform normal vectors in a shader, the lighting is wrong. When a non-uniform scale is involved (for example, shrinking only x by half), a vector that was perpendicular to the surface is no longer perpendicular. The correct answer is to use the transpose of the inverse of the model matrix, and this fact does not stick unless you have worked with matrices yourself.

Another is floating-point accumulation. If you multiply a rotation matrix a little every frame and accumulate it, after a few hundred frames orthogonality breaks down and the object slowly gets distorted. That is why engines keep an angle or a quaternion as the state and build the matrix anew every frame. They do not keep the matrix itself as state.

Another thing often missed is the cost of the inverse matrix. A general 4x4 inverse is heavy to compute, but most transformations handled in graphics consist of only rotation and translation, so they can be inverted much more cheaply. The rotation part is an orthogonal matrix, so its transpose is its inverse, and for the translation you flip the sign and apply the inverse of the rotation. You actually use this property when building a camera matrix — the view matrix we saw earlier, which stands the camera's three axes as rows and puts negative dot products in the last column, is exactly this computation.

Finally, you also need a feel for choosing which vector operation to use. The dot product is for angles and projection, and the cross product is for perpendicular directions and area. If you want to know whether two vectors face the same direction, look at the sign of the dot product, and if you want to know which way they turn, look at the sign of the cross product. These two sentences explain a good part of graphics code.

What you will do in the next lab

You build 3D vector operations and 4x4 matrix multiplication yourself, and assemble translation, scale and rotation matrices. After checking in numbers how the result changes when you swap the multiplication order, you draw three transformed squares on one PNG and look at them. Finally, you find the normal and area of a triangle with the cross product.