A Shader Is Called Twice
In one line
The vertex shader is called once per vertex and produces clip coordinates and the values to interpolate, and the fragment shader is called once per pixel and produces a color. Between them, the rasterizer blends the values with barycentric coordinates.
Why this was needed
Old graphics hardware was fixed-function. The lighting model and the way textures were combined were baked into the hardware, so developers could only turn the knobs provided. To express something new, you had to wait for the hardware to change.
Shaders opened up those two points to be programmable. Where to send a vertex and what color to paint a pixel are decided by code the developer writes. What matters is that it was not opened just anywhere. The structure of the pipeline itself was left as is, and only the vertex stage and the fragment stage were opened. That lets the hardware keep processing the rest (rasterization, depth testing, blending) in parallel.
How it works
The vertex shader takes one vertex and produces one vertex. There are two kinds of input.
attribute 정점마다 다른 값 — 위치, 법선, UV, 정점 색
uniform 드로우 콜 전체에서 같은 값 — MVP 행렬, 광원 방향, 시간
The output must include the clip coordinates (gl_Position), and besides that it sends the values to pass to the fragment stage out as varying (in today's naming, out). A vertex shader cannot create or remove vertices. One goes in and one comes out. That lets the hardware know in advance how many to process and split them up in parallel.
The rasterizer is a fixed stage that cannot be programmed. It takes a triangle, picks out the pixels inside it, blends the varyings of the three vertices at each pixel with barycentric coordinates, and hands them to the fragment shader. What you built yourself in the previous course is exactly this stage.
The fragment shader takes one pixel and produces one color. Its inputs are the interpolated varyings, uniforms, and textures. There is one special thing you can do here, and that is discard. If you declare that this fragment will not be used at all, nothing is left in either the depth buffer or the color buffer. It is used to express surfaces with holes, like leaves or wire mesh.
One point to note is that a fragment is different from a pixel. A fragment is "a candidate that some triangle produced at this pixel position", and several fragments can arise at the same pixel. Which of them becomes the final pixel is decided by depth testing and blending.
These two are not the only stages. Between vertex and fragment you can insert a geometry shader and tessellation, and these can increase or decrease vertices. In exchange, the count is decided at run time, so the hardware cannot divide the work in advance and they are much slower. So in practice they are used only where truly necessary. Outside the pipeline there is also the compute shader, which is a channel for running arbitrary computation on the GPU regardless of rasterization.
One more thing to know is the cost of passing uniforms and attributes. A uniform only needs to be set once per draw call, but the draw call itself is expensive, so when drawing a thousand objects with the same shader, changing the uniform a thousand times and drawing a thousand times is a bad design. Instancing is a method of passing the value that differs per object like an attribute and handling it all in one draw call, and scenes that draw tens of thousands of trees or blades of grass are built that way.
What it looks like in the field
This is why, when looking at performance, you count the number of vertices and the number of fragments separately. The vertex shader is called as many times as there are vertices, and the fragment shader is called as many times as the area covering the screen. One triangle that fills the screen is three vertices and hundreds of thousands of fragments. So you move heavy computation to the vertex shader where possible and leave it to interpolation. However, interpolation is linear, so a value that needs normalization (like a normal) must be normalized again on the fragment side.
Another is discard is not free. A shader with discard cannot use the optimization that tests depth in advance. This is because you can only know whether this fragment survives after computing the color. So a surface with holes is more expensive than a completely opaque one.
And the contract between shaders is made of names and types. If the vertex shader sends out out vec2 uv, the fragment shader receives it as in vec2 uv. If these names do not match, the connection breaks, and in some cases compilation passes and only the value comes in as 0, which is hard to find. Modern GLSL provides a way to pin the location with a number (layout(location = 0)) to reduce this problem.
Another thing to know is the cost of branching. A GPU executes the same instruction together for a group of dozens of fragments. If the result of an if splits within the group, it executes both branches and throws away the unneeded results, so even with a conditional it does not get faster and actually costs the price of running both. If the whole group goes the same way, then it really does skip. So it is better to design shader conditionals to split in large contiguous regions of the screen.
What you will do in the next lab
You write a vertex shader and a fragment shader as Python functions, place a rasterizer between them, and draw one triangle. You confirm that changing only a uniform gives a different picture from the same geometry, and you punch holes with discard. Finally, you count the number of vertices, triangles and fragments and see in numbers how many times each stage was called.