3D Math and a Software Rasterizer
Why an Image File Looks the Way It Does
In one line
An image is ultimately just an array of numbers, and a file format is nothing more than an agreement on the order in which that array is laid out and how it is shrunk. PPM chose not to shrink it, and PNG chose to shrink it.
Why this was needed
The first wall people hit when learning 3D graphics is not the math but having no way to see the result. Opening a window requires a window manager, and saving a picture requires an image library. If the lab environment has neither, people give up on learning.
Yet writing an image file yourself is far simpler than it sounds, and you get something out of the process. If you skip ahead with a single line like image.save("a.png"), you never find out in what order pixels sit in memory, why odd widths cause alignment problems, or why an alpha channel makes a file four thirds as large. In graphics, these things come back later as performance problems.
How it works
PPM (Portable PixMap) is a format from the Netpbm family that lists pixels as they are after a header. This is the entire structure of P6, the binary variant.
P6 <- 매직 넘버 (P3 이면 픽셀을 십진수 글자로 적는다)
256 256 <- 폭과 높이
255 <- 채널 최댓값
<픽셀 바이트> <- R,G,B,R,G,B, ... 왼쪽 위부터 오른쪽으로, 그다음 아래 줄로
Pixels are laid out in row-major order. The top-left pixel comes first, and after one row is filled to the end you move down to the next row. So the byte position of coordinate (x, y) is (y * 폭 + x) * 3 (the placeholder is the image width). That y grows downward is something the format decided, and it clashes with 3D math, where y grows upward. Where to flip this mismatch is a problem that keeps coming up later.
The trouble with PPM is that browsers cannot display it. That is why PNG is needed. A PNG file starts with an 8-byte signature, followed by chunks made up of length, type, content and CRC. The minimum set of chunks you need is three.
89 50 4E 47 0D 0A 1A 0A <- 서명
IHDR 폭, 높이, 비트깊이 8, 색 타입 2(트루컬러)
IDAT 각 행 앞에 필터 바이트 1개를 붙인 뒤 zlib 으로 압축한 것
IEND 끝
The key point is that one filter byte is attached in front of every row. A filter is a device that raises the compression ratio by storing only the difference from neighboring pixels, and filter 0 means "no filter". If you use only filter 0, the encoder simply puts one 0 in front of each row and is done. The Python standard library zlib does both the compression and the CRC for you — zlib.compress and zlib.crc32. That is why a PNG encoder fits in twenty lines.
Adding an alpha channel makes the color type 6 and raises the bytes per pixel from 3 to 4. That is where the file growing to four thirds of its size comes from. This one channel often becomes a problem when you calculate texture memory: if you promote even textures that do not actually use alpha to four channels, you need 33% more GPU memory.
Row spacing is also worth knowing. PNG and PPM pack rows tightly at 폭 × 채널 수 bytes (the placeholders are the width and the number of channels), but a graphics API framebuffer often aligns each row to 4 bytes or a multiple of 4. In that case the actual row pitch (stride) is larger than the value computed from the width. If you read while ignoring this difference, the rows shift a little at a time and the picture slants diagonally.
What it looks like in the field
There are two bugs people hit first when building a renderer, and both come from this layer. One is the picture coming out flipped upside down. The OpenGL framebuffer has its origin at the bottom left, while image formats have it at the top left, so if you do not flip when reading and saving, the picture comes out upside down as is. The other is the picture coming out shifted diagonally, which is almost always a wrong row length (stride) calculation. It happens when width × number of channels differs from the actual row pitch.
It is also useful for debugging. When a shader result looks wrong, the fastest diagnosis is to paint the intermediate values (normal, depth, UV) directly as colors and dump them as an image. For that, you need a function at hand that "drops the array of numbers you currently have into a file".
The meaning of color values also deserves a mention. The 0–255 written in a file is usually a gamma-corrected value, not a value proportional to light intensity. If you average two integers to find the midpoint of two colors, the result is not the midpoint in actual light intensity. You can ignore this in this lab, but when lighting calculations later look somehow muddy, this is the place to suspect.
What you will do in the next lab
You write a PPM by hand to make a 4×4 picture, widen it into a 256×256 gradient, and then write the same pixels as a PNG using only zlib and struct. After that you start a small HTTP server inside the Pod and view your own picture through the web preview. Finally you draw a circle and lines, and check that the values you count from the picture match the values you wrote down.