TT Lab
Get started
Learn Learning paths Courses

3D Math and a Software Rasterizer

Stand a Cube Up on the Screen

Continue in TT Lab

Goal

Build the camera matrix and the projection matrix yourself, and stand a cube up on the screen through the perspective divide and the viewport transform. When this lab is done, you can explain "what a vertex shader does" in one line.

Why it matters

A single 3D point passes through four coordinate systems before it becomes a pixel coordinate: model, world, view and clip. Each step is one matrix, so you can multiply them in advance into an MVP, and that is why the GPU does just one multiplication per vertex.

The most important thing here is the fact that perspective comes from a division, not from a matrix. Matrix multiplication is a linear operation and cannot express a division. So the projection matrix sets its last row to (0,0,-1,0) to load the distance from the camera into the result's w, and then divides x, y and z by that w. If you understand this detour, it also explains why the precision of the depth buffer is concentrated close to the camera and why you must not shrink the near plane carelessly.

Steps

  1. Put the toolbox gfxlib.py and mat3d.py in /root/pipeline.
  2. Make look_at in /root/pipeline/view.py.
  3. Make perspective in /root/pipeline/proj.py.
  4. Make to_screen in /root/pipeline/screen.py and write /root/pipeline/out/04-divide.txt.
  5. Draw a cube wireframe to /root/pipeline/out/cube.png.
  6. Add orthographic and make /root/pipeline/out/ortho.png.
  7. Draw with three fields of view and write /root/pipeline/out/07-fov.txt.

Notes

Put the drawing toolbox in place

Save /root/pipeline/gfxlib.py exactly as in the example, and use /root/pipeline/check.py to draw a test pattern and make /root/pipeline/out/00-check.png. The pattern is a 64x64 black background with a white (255,255,255) diagonal line from (0,0) to (63,63), and over it a red (255,0,0) horizontal line from (0,32) to (63,32). This lab also uses the matrix tool /root/pipeline/mat3d.py.

From this lab on, you do not rebuild the PNG encoder. We hand you the same code you built by hand in the first lab as a tool — because file formats are not what you learn here.

The lab Pod has no volume, so the files you made in the previous lab are not kept. That is why each lab starts by putting the toolbox in place again.

Create Canvas(w, h, bg), draw the two lines with line(x0, y0, x1, y1, rgb), and then save with write_png(path). Draw the horizontal line later, so that the intersection (32,32) becomes red.

mat3d.py is the same as the vector and matrix operations you built yourself in the previous lab. Here you receive it as material to use, and what you learn is the camera and the projection.

Pull the camera to the origin

In /root/pipeline/view.py, make look_at(eye, target, up) return a 4x4 view matrix. The three axes are f = normalize(target-eye), s = normalize(cross(f, up)) and u = cross(s, f), and you put -f in the third row.

Each row of the matrix is a camera axis, and the last column is the value that moves the camera position back along that axis direction. That is, row0 = (s.x, s.y, s.z, -dot(s, eye)).

The reason only the third row has the opposite sign is that, in a right-handed coordinate system, we agreed that the camera looks in the -z direction. So that row becomes (-f.x, -f.y, -f.z, dot(f, eye)).

Checking is simple. With eye=(0,0,5), target=(0,0,0) and up=(0,1,0), the origin must go to (0, 0, -5). That means a spot 5 in front of the camera.

For vector operations, use normalize, cross, dot and vsub from mat3d.

Squeeze the frustum into a cube

In /root/pipeline/proj.py, make perspective(fov_deg, aspect, near, far). Let t = 1/tan(fov/2), and set m[0][0]=t/aspect, m[1][1]=t, m[2][2]=-(far+near)/(far-near), m[2][3]=-2*far*near/(far-near) and m[3][2]=-1, with everything else 0.

You could say that the last row being (0, 0, -1, 0) is the whole point of this matrix. Because of this row, the w component of the result becomes the distance from the camera (-z), and the division in the next step creates the perspective.

m[3][3] is 0. If you set it to 1, w no longer depends on distance and the perspective disappears.

It is t = 1/tan(radians(fov)/2). If you forget to convert fov to radians, the field of view comes out wrong, and no error appears.

Divide by w and get screen coordinates

In /root/pipeline/screen.py, make to_screen(clip, w, h). It takes clip coordinates (x,y,z,w), divides to get ndc = (x/w, y/w), and returns two real numbers, sx = (ndc_x*0.5+0.5)*w_px and sy = (1-(ndc_y*0.5+0.5))*h_px. Then, using perspective(60,1,0.1,100), map the view-space points A(1,0,-2) and B(1,0,-8) onto a 256x256 screen and write the values to /root/pipeline/out/04-divide.txt as two lines, A=sx,sy and B=sx,sy (six decimal places).

The two points have the same x of 1 and differ only in depth. If you did the division correctly, B, the farther one, must come out closer to the screen center (128). That is perspective.

Subtracting from 1 in sy is the y flip. This is because y in an image grows downward.

mat3d.apply(P, (1,0,-2)) multiplies with w set to 1 and returns four components. Put that result into to_screen.

The twelve edges of a cube

Use /root/pipeline/cube.py to make /root/pipeline/out/cube.png (256x256, black background, white lines). The cube is eight vertices (±1,±1,±1) and the twelve edges joining them, the camera is look_at((2.5,2.5,2.5),(0,0,0),(0,1,0)), and the projection is perspective(60, 1, 0.1, 100).

An edge is a pair of vertices whose coordinates differ in only one axis. If you pick such pairs out of the eight, you get exactly twelve.

The MVP is mat3d.mul(P, V). The model matrix is the identity matrix, so it is omitted, and in the multiplication order the projection is on the left.

Transform each vertex only once to build a list of screen coordinates, and join the edges with cv.line by taking points from that list. If you recompute for every edge, you do the same work twice.

Compare with orthographic projection

Add orthographic(l, r, b, t, n, f) to /root/pipeline/proj.py, and make /root/pipeline/out/ortho.png with the same camera. The bounds are l=-2.5, r=2.5, b=-2.5, t=2.5, n=0.1, f=100.

The last row of the orthographic projection matrix is (0, 0, 0, 1). Since w is always 1, the division becomes the identity, and therefore size is the same regardless of distance.

The diagonal entries are 2/(r-l), 2/(t-b) and -2/(f-n), and the last column is -(r+l)/(r-l), -(t+b)/(t-b) and -(f+n)/(f-n). It combines a translation and a scale that map the range to -1 to 1.

If you put the two pictures side by side, you can see the difference. In perspective, the near face is large and the far face is small, but in orthographic projection the two opposite faces are the same size, so it looks like a hexagon.

What changes when you change the field of view

Draw the same cube with fields of view of 60, 75 and 90 degrees to make /root/pipeline/out/fov60.png, fov75.png and fov90.png, and write the horizontal width the cube occupies in each picture (the maximum minus the minimum of the screen x coordinates of the eight vertices) to /root/pipeline/out/07-fov.txt as three lines, fov60=, fov75= and fov90= (three decimal places).

Leave the camera position as it is and change only the field of view. When the field of view widens, more comes into the screen, so the same object actually gets smaller. The three values must come out in decreasing order.

Do not measure the width in the picture; compute it directly from the transformed vertex coordinates. Subtract the minimum from the maximum of the eight sx values.

If you wrap the code from the previous step in a function that takes only the field of view as an argument, you do not have to repeat it three times.