Compilers — Build a Small Language from Start to Finish
Registers, Frames and 16 Bytes — The Calling Convention
In one line
Native code generation is translating the tree into CPU instructions (here, x86-64 in GNU as syntax). The value of an expression is kept in a register (%rax), put briefly on the stack when registers run short, and functions pass arguments in registers and set up frames as laid down by the calling convention (System V AMD64). Of those promises, the one most quietly broken is the rule to align the stack to 16 bytes just before a call.
Why this was needed
In module 6's VM, a Python loop turns once for every instruction it executes. Native code has no such loop — the CPU reads the instructions directly. In exchange, the compiler has to decide everything the VM used to do for it: which register to put a value in, where in memory to put local variables, and, when calling a function, where to pass the arguments and how to keep the return address safe.
Some of these decisions are not ours to make freely. For the code we produce to call the C library (printf) or code produced by another compiler, and vice versa, we must keep the same promises. On Linux x86-64, those promises are the System V AMD64 ABI.
How it works
Expressions in %rax, on the stack when short. This course's code generator uses the simplest approach. The value of every expression is kept in %rax. For a binary operation, it computes the left side and puts it on the stack with push %rax, computes the right side and moves it to %rcx, and then pops the left side back with pop %rax and operates. It does not use ways of saving registers (register allocation in module 9), so it is slow, but it can translate any expression without going wrong.
print 1 + 2 * 3; main 의 프레임
movabs $1, %rax 높은 주소
push %rax ← 왼쪽(1)을 올림 │ 돌아갈 주소 │ call 이 넣음
movabs $2, %rax │ 이전 %rbp │ ← %rbp
push %rax ← 왼쪽(2)을 올림 │ 지역 변수 슬롯 0 │ -8(%rbp)
movabs $3, %rax │ 지역 변수 슬롯 1 │ -16(%rbp)
mov %rax, %rcx │ (16의 배수로 맞춤) │
pop %rax ← 2 │ 식 계산용 push … │ ← %rsp
imul %rcx, %rax ← 6 낮은 주소
mov %rax, %rcx
pop %rax ← 1
add %rcx, %rax ← 7
mov %rax, %rdi ← 첫 인자
call mini_print_int
The calling convention. Integer arguments go in the order %rdi %rsi %rdx %rcx %r8 %r9, and the return value in %rax. The receiving side sets up a frame with push %rbp · mov %rsp, %rbp · sub $N, %rsp, and moves the arguments into its own slots (-8(%rbp) …) — because those registers are used by other calls while the body is being computed. When it finishes, it clears the frame and returns with leave · ret. If you compute the arguments from the left and push them in turn, and after computing them all pop them in reverse order into the registers, it is safe even if a call during argument computation overwrites the registers.
16-byte alignment. The ABI requires %rsp to be a multiple of 16 at the moment of call. On entering a function, it is off by 8 because of the return address (8 bytes), push %rbp makes it aligned again, and if you take the frame size as a multiple of 16, it is aligned at the start of the body. After that, every time one more push for expression computation is added, it alternates between off by 8 and aligned. So if you call while the number of slots you have pushed is odd, you break the rule — this is the case where you call on the right side while the left side is pushed, as in 1 + f(2). In that case, align with sub $8, %rsp, call, and after returning restore with add $8, %rsp.
Breaking this rule is usually quiet. It dies with a segmentation fault only on the day it meets an SSE instruction that requires 16-byte alignment (movaps) inside printf — it is the worst kind of bug, one that works for some inputs and not for others. This lab's runtime helper (/opt/fixtures/mini/runtime.c) checks the alignment every time it is called and tells you right away.
Runtime helpers. Rather than having the code generator scatter checks for division by zero, INT64_MIN / -1, and negative exponents around each expression, it hands them along with the line and column to helpers written in C (mini_div, mini_mod, and mini_pow). idiv raises a CPU exception (SIGFPE) on INT64_MIN / -1, so using it as it is would make the meaning diverge from the interpreter.
What it looks like in the field
- The shape of
gcc -O0 -S. The output of gcc without optimization resembles this approach too — every local variable is in a stack slot like-8(%rbp), and it reads and writes each time it is used. If you look at it side by side with the-O2output in module 10, you can see what register allocation and optimization strip away. - Alignment incidents. It is actually common for hand-written assembly or JIT-generated code to break alignment when calling a C function and die at
movaps. Stack alignment shows up in bug reports as "a segmentation fault inside printf only for certain inputs." - Other ABIs. On Windows x64, the argument registers are
%rcx %rdx %r8 %r9and the caller must set aside 32 bytes of "shadow space." That the promises differ per operating system even on the same CPU is a point often hit in cross-compilation.
What you will do in the next lab
In codegen.py, you build in turn main and output, arithmetic and comparisons and runtime helper calls, global variables (.data), block local variables (frame slots), if, while, and short-circuit evaluation (labels and jumps), and functions and argument registers; you check that stack alignment is kept for calls in the middle of a computation, and then build build, which translates the whole program to assembly and bakes it with gcc. At every step, the grader actually bakes and runs your assembly and checks whether it prints the same lines as the interpreter.