3D Math and a Software Rasterizer
Stand a Cube Up on the Screen
Goal
Build the camera matrix and the projection matrix yourself, and stand a cube up on the screen through the perspective divide and the viewport transform. When this lab is done, you can explain "what a vertex shader does" in one line.
Why it matters
A single 3D point passes through four coordinate systems before it becomes a pixel coordinate: model, world, view and clip. Each step is one matrix, so you can multiply them in advance into an MVP, and that is why the GPU does just one multiplication per vertex.
The most important thing here is the fact that perspective comes from a division, not from a matrix. Matrix multiplication is a linear operation and cannot express a division. So the projection matrix sets its last row to (0,0,-1,0) to load the distance from the camera into the result's w, and then divides x, y and z by that w. If you understand this detour, it also explains why the precision of the depth buffer is concentrated close to the camera and why you must not shrink the near plane carelessly.
Steps
- Put the toolbox
gfxlib.pyandmat3d.pyin/root/pipeline. - Make
look_atin/root/pipeline/view.py. - Make
perspectivein/root/pipeline/proj.py. - Make
to_screenin/root/pipeline/screen.pyand write/root/pipeline/out/04-divide.txt. - Draw a cube wireframe to
/root/pipeline/out/cube.png. - Add
orthographicand make/root/pipeline/out/ortho.png. - Draw with three fields of view and write
/root/pipeline/out/07-fov.txt.
Notes
- To see the pictures, start a server with
nohup python3 -m http.server 8080 -d /root/pipeline/out &and then openhttp://localhost:8080/in the Web preview. If you put cube.png and ortho.png side by side, the difference is clear. - Common mistake 1: setting
m[3][3]of the projection matrix to 1. Then w no longer depends on distance and the perspective disappears. - Common mistake 2: leaving out the minus sign in the third row of
look_at. Then the camera looks the opposite way and the screen is empty.
Put the drawing toolbox in place
Save /root/pipeline/gfxlib.py exactly as in the example, and use /root/pipeline/check.py to draw a test pattern and make /root/pipeline/out/00-check.png. The pattern is a 64x64 black background with a white (255,255,255) diagonal line from (0,0) to (63,63), and over it a red (255,0,0) horizontal line from (0,32) to (63,32). This lab also uses the matrix tool /root/pipeline/mat3d.py.
From this lab on, you do not rebuild the PNG encoder. We hand you the same code you built by hand in the first lab as a tool — because file formats are not what you learn here.
The lab Pod has no volume, so the files you made in the previous lab are not kept. That is why each lab starts by putting the toolbox in place again.
Create Canvas(w, h, bg), draw the two lines with line(x0, y0, x1, y1, rgb), and then save with write_png(path). Draw the horizontal line later, so that the intersection (32,32) becomes red.
mat3d.py is the same as the vector and matrix operations you built yourself in the previous lab. Here you receive it as material to use, and what you learn is the camera and the projection.
Pull the camera to the origin
In /root/pipeline/view.py, make look_at(eye, target, up) return a 4x4 view matrix. The three axes are f = normalize(target-eye), s = normalize(cross(f, up)) and u = cross(s, f), and you put -f in the third row.
Each row of the matrix is a camera axis, and the last column is the value that moves the camera position back along that axis direction. That is, row0 = (s.x, s.y, s.z, -dot(s, eye)).
The reason only the third row has the opposite sign is that, in a right-handed coordinate system, we agreed that the camera looks in the -z direction. So that row becomes (-f.x, -f.y, -f.z, dot(f, eye)).
Checking is simple. With eye=(0,0,5), target=(0,0,0) and up=(0,1,0), the origin must go to (0, 0, -5). That means a spot 5 in front of the camera.
For vector operations, use normalize, cross, dot and vsub from mat3d.
Squeeze the frustum into a cube
In /root/pipeline/proj.py, make perspective(fov_deg, aspect, near, far). Let t = 1/tan(fov/2), and set m[0][0]=t/aspect, m[1][1]=t, m[2][2]=-(far+near)/(far-near), m[2][3]=-2*far*near/(far-near) and m[3][2]=-1, with everything else 0.
You could say that the last row being (0, 0, -1, 0) is the whole point of this matrix. Because of this row, the w component of the result becomes the distance from the camera (-z), and the division in the next step creates the perspective.
m[3][3] is 0. If you set it to 1, w no longer depends on distance and the perspective disappears.
It is t = 1/tan(radians(fov)/2). If you forget to convert fov to radians, the field of view comes out wrong, and no error appears.
Divide by w and get screen coordinates
In /root/pipeline/screen.py, make to_screen(clip, w, h). It takes clip coordinates (x,y,z,w), divides to get ndc = (x/w, y/w), and returns two real numbers, sx = (ndc_x*0.5+0.5)*w_px and sy = (1-(ndc_y*0.5+0.5))*h_px. Then, using perspective(60,1,0.1,100), map the view-space points A(1,0,-2) and B(1,0,-8) onto a 256x256 screen and write the values to /root/pipeline/out/04-divide.txt as two lines, A=sx,sy and B=sx,sy (six decimal places).
The two points have the same x of 1 and differ only in depth. If you did the division correctly, B, the farther one, must come out closer to the screen center (128). That is perspective.
Subtracting from 1 in sy is the y flip. This is because y in an image grows downward.
mat3d.apply(P, (1,0,-2)) multiplies with w set to 1 and returns four components. Put that result into to_screen.
The twelve edges of a cube
Use /root/pipeline/cube.py to make /root/pipeline/out/cube.png (256x256, black background, white lines). The cube is eight vertices (±1,±1,±1) and the twelve edges joining them, the camera is look_at((2.5,2.5,2.5),(0,0,0),(0,1,0)), and the projection is perspective(60, 1, 0.1, 100).
An edge is a pair of vertices whose coordinates differ in only one axis. If you pick such pairs out of the eight, you get exactly twelve.
The MVP is mat3d.mul(P, V). The model matrix is the identity matrix, so it is omitted, and in the multiplication order the projection is on the left.
Transform each vertex only once to build a list of screen coordinates, and join the edges with cv.line by taking points from that list. If you recompute for every edge, you do the same work twice.
Compare with orthographic projection
Add orthographic(l, r, b, t, n, f) to /root/pipeline/proj.py, and make /root/pipeline/out/ortho.png with the same camera. The bounds are l=-2.5, r=2.5, b=-2.5, t=2.5, n=0.1, f=100.
The last row of the orthographic projection matrix is (0, 0, 0, 1). Since w is always 1, the division becomes the identity, and therefore size is the same regardless of distance.
The diagonal entries are 2/(r-l), 2/(t-b) and -2/(f-n), and the last column is -(r+l)/(r-l), -(t+b)/(t-b) and -(f+n)/(f-n). It combines a translation and a scale that map the range to -1 to 1.
If you put the two pictures side by side, you can see the difference. In perspective, the near face is large and the far face is small, but in orthographic projection the two opposite faces are the same size, so it looks like a hexagon.
What changes when you change the field of view
Draw the same cube with fields of view of 60, 75 and 90 degrees to make /root/pipeline/out/fov60.png, fov75.png and fov90.png, and write the horizontal width the cube occupies in each picture (the maximum minus the minimum of the screen x coordinates of the eight vertices) to /root/pipeline/out/07-fov.txt as three lines, fov60=, fov75= and fov90= (three decimal places).
Leave the camera position as it is and change only the field of view. When the field of view widens, more comes into the screen, so the same object actually gets smaller. The three values must come out in decreasing order.
Do not measure the width in the picture; compute it directly from the transformed vertex coordinates. Subtract the minimum from the maximum of the eight sx values.
If you wrap the code from the previous step in a function that takes only the field of view as an argument, you do not have to repeat it three times.