TT Lab
Get started
Learn Learning paths Courses

Shaders and the GPU Pipeline

Two Shaders and What Lies Between

Continue in TT Lab

Goal

Write a vertex shader and a fragment shader as Python functions, place a rasterizer between them, and actually produce a picture. When this lab is done, you can see where each value comes from and where it goes when you read GLSL code.

Why it matters

Shaders are not opened at just any point in the pipeline. The structure was left as is, and only the vertex stage and the fragment stage were opened. That lets the hardware keep processing the rest (rasterization, depth testing, blending) in parallel.

That constraint creates the rules of shaders. A vertex shader cannot create or remove vertices, and a fragment shader cannot see the results of neighboring pixels. In exchange, the channel between the two stages is clear — when the vertex shader sends out a varying, the rasterizer blends it with barycentric coordinates and hands it to the fragment shader. Once you have built this structure by hand, you understand why the constraints of GPU programming take the shape they do.

Steps

  1. Put the toolbox in /root/stages.
  2. /root/stages/vs.py — the vertex shader.
  3. /root/stages/varying.py — value interpolation.
  4. /root/stages/fs.py and pipeline.py — the fragment shader and out/uv.png.
  5. /root/stages/out/u1.png, u2.png, u3.png — changing only the uniform.
  6. /root/stages/out/discard.png — discarding fragments.
  7. /root/stages/out/final.png and out/07-counts.txt — the call counts.

Notes

Put the drawing toolbox in place

Save /root/stages/gfxlib.py exactly as in the example, and use /root/stages/check.py to draw a test pattern and make /root/stages/out/00-check.png. The pattern is a 64x64 black background with a white (255,255,255) diagonal line from (0,0) to (63,63), and over it a red (255,0,0) horizontal line from (0,32) to (63,32).

From this lab on, you do not rebuild the PNG encoder. We hand you the same code you built by hand in the first lab as a tool — because file formats are not what you learn here.

The lab Pod has no volume, so the files you made in the previous lab are not kept. That is why each lab starts by putting the toolbox in place again.

Create Canvas(w, h, bg), draw the two lines with line(x0, y0, x1, y1, rgb), and then save with write_png(path). Draw the horizontal line later, so that the intersection (32,32) becomes red.

In this lab you use it to output the pixels the shaders produce as pictures.

The vertex shader

In /root/stages/vs.py, make vertex_shader(attr, uniforms). attr is {"position": (x,y,z), "uv": (u,v)} and uniforms is {"mvp": 4x4 행렬, "scale": 실수} (the placeholders are a 4x4 matrix and a real number). Multiply the position by scale and then apply mvp, and return {"gl_Position": ..., "varyings": {...}}, where the four components are the gl_Position and a dictionary that holds uv as it is is the varyings.

A vertex shader takes one vertex and produces one vertex. It cannot create or remove them. That lets the hardware know in advance how many to process and split them up in parallel.

An attribute is a value that differs per vertex, and a uniform is a value that is the same across the whole draw call. This distinction decides where the GPU puts the values.

For the matrix multiplication, compute m[r][0]*x + m[r][1]*y + m[r][2]*z + m[r][3]*1 for r=0..3. It is a point, so w is 1.

varyings are the values to pass to the fragment stage. Here there is only uv, but in a real shader normals, colors, tangents and so on are carried together.

The rasterizer blends the values

In /root/stages/varying.py, make interpolate(bary, v0, v1, v2). bary is the three real numbers (u, v, w), and v0, v1 and v2 are varyings dictionaries with the same keys. If a value is a tuple, compute u*v0 + v*v1 + w*v2 per component, and if it is a real number, compute it as is, and return a dictionary of the same shape.

This step is the job not of a shader but of the rasterizer. It is a fixed stage that cannot be programmed, and the reason we build it ourselves is that this is the channel through which the vertex shader's output becomes the fragment shader's input.

You must tell the kinds of values apart. Split with isinstance(x, (tuple, list)).

The barycentric coordinates sum to 1, so if the values at the three vertices are all the same, the result is the same value too. You can check your implementation with this property.

The fragment shader and the first picture

In /root/stages/fs.py, make fragment_shader(varyings, uniforms). Set u, v = varyings["uv"] and k = uniforms["brightness"] and return (255*(0.2+0.8*u)*k, 255*(0.2+0.8*v)*k, 255*(0.2+0.8*(1-u))*k). Then, with /root/stages/pipeline.py, draw the three vertices (-0.8,-0.8,0)/uv(0,0), (0.8,-0.8,0)/uv(1,0) and (-0.8,0.8,0)/uv(0,1) with an identity-matrix mvp, scale 1 and brightness 1 to make /root/stages/out/uv.png (256x256, black background). The screen coordinates are sx=(x*0.5+0.5)*256 and sy=(1-(y*0.5+0.5))*256.

The order in the pipeline is this. Call the vertex shader for each vertex, move the clip coordinates to screen coordinates, scan the triangle's bounding box finding the barycentric coordinates, and if the pixel is inside, interpolate the varyings and call the fragment shader.

The fragment shader is called per pixel. This triangle covers nearly one third of the screen, so it is called more than 20,000 times. The vertex shader is called only three times. This difference is the starting point of any talk about performance.

The reason 0.2 is added in the color calculation is to keep the color from becoming 0 even where uv is 0. It is needed when you compare while changing the brightness in the next step.

The w of gl_Position is 1, so the division is the identity here. When perspective is involved, that division comes alive.

Change only the uniform

Leave the geometry as it is and change only brightness to 1.0, 0.6 and 0.3 to make three pictures, /root/stages/out/u1.png, u2.png and u3.png.

The vertex data and the triangle stay as they are. Only one uniform changes.

In the three pictures, the painted positions (the set of non-black pixels) must be exactly the same. After all, the geometry did not change. If they differ, the vertex shader or the rasterizer is being affected by the uniform.

The colors, on the other hand, must differ noticeably. This is what uniforms are for — when you draw the same mesh several times with a different material or different lighting, you change only the uniform and draw, without sending the vertex data again.

Because 0.2 is added in the color calculation, the pixels do not turn black even at brightness 0.3.

Discard fragments

Draw the same triangle, but discard the fragments whose interpolated uv satisfies u*u + v*v > 0.36, to make /root/stages/out/discard.png (brightness is 1.0).

Discarding is something only the fragment shader can do. If you declare that this fragment will not be used at all, nothing is left in either the color buffer or the depth buffer.

It is used to express surfaces with holes, like leaves or wire mesh. It is different from translucency — translucency blends so that what is behind shows through, and discarding draws nothing at that spot.

Discarding is not free. A shader with discard can only know whether a fragment survives after computing the color, so it makes the optimization that tests depth in advance unusable. That is why a surface with holes is more expensive than a completely opaque one.

The discarded spots keep the background (black) as it is.

Which stage was called how many times

Draw again with discarding on to make /root/stages/out/final.png, and write four lines to /root/stages/out/07-counts.txt: vertices=, triangles=, fragments_in= and fragments_out=. fragments_in is the number of fragments that came inside the triangle, and fragments_out is the number left after discarding.

The vertex shader is called as many times as there are vertices, and the fragment shader is called as many times as the area covering the screen. This triangle has three vertices, yet there are more than 20,000 fragments.

So you move heavy computation to the vertex shader where possible and leave it to interpolation. However, interpolation is linear, so a value that needs normalization must be normalized again on the fragment side.

fragments_out must be exactly equal to the number of non-black pixels in the final picture. The grader opens the picture, counts, and compares. If they differ, one of the two is wrong.

Do not make up the numbers; count them inside the loop.