TT Lab
Get started
Learn Learning paths Courses

My TCP Parcel Arrived in Pieces

Stop Giving Slow Parcels Unlimited Time

Continue in TT Lab

In one line

Length and time need different limits. Making a little progress each time does not mean you can wait forever.

Why this was needed

The parcel envelope says that a body of nearly 4 GiB is coming. The value itself fits in a 32-bit integer, but our service's receive budget is 4096 bytes. If you trust this number and allocate memory first, a single small header consumes a large resource. The other side's promise is not permission to allocate resources. This lab is practice in rejecting a large body at the moment the header is complete, without waiting for the body to actually arrive.

After you limit memory, time remains. If you wait a fresh 1 second for every byte that arrives, the connection stays alive as long as the other side sends one byte every 0.9 seconds. The requirement that the whole frame be received within 1 second is different from the requirement that each recv return within 1 second. For the former, you must pass the remaining budget all the way to the end of the request.

How it works

At the start, calculate deadline=clock()+timeout once. Right before each recv call, calculate remaining=deadline-clock(). If remaining is 0 or less, raise TimeoutError before reading any more. Otherwise, use sock.settimeout(remaining) so that even a blocking read does not exceed the remaining time. The header and the body must share the same deadline. You do not grant new time for the body just because the header arrived late.

The default clock is time.monotonic. It keeps elapsed-time calculations from running backward even if the wall clock moves back because of a clock adjustment. In tests, you inject a function through the clock argument. Then, instead of actually sleeping for several seconds to wait for edge cases, you can advance a fake time a little at a time as recv progresses. A fake-time test checks the time budget calculation deterministically, and a real socket test checks the coupling with the operating system interface. You do not claim that one of them alone proves the other.

read_exact repeats recv for the remaining size until it has n bytes. It is normal for recv(n) to return fewer than n bytes. b"" is EOF. If EOF arrives without any new header received at all, there are no more messages, so it returns None. If even one header byte was received or part of the body is missing, it is an EOFError. A body of length 0 has no bytes to read, so it returns b"" right away. If you call recv again here, you may eat the next frame or wait for nothing.

You must restore the socket timeout from before the call not only on success but also when an exception occurs. If you hand the same socket to the next task while the short remaining value from the last read is still set, the next task fails early for no apparent reason. try/finally expresses this resource contract. If timeout is 0, negative, infinite, or NaN, it violates the API contract of a positive finite budget, so it is a ValueError.

After a length error, discard that connection. This protocol does not define a recovery that discards bytes one at a time and looks for the next header. So if you resynchronize because the next number looks plausible, you mistake the middle of a body for a new message. Whether a real protocol defines recovery points or checksums is a separate question. The parser does not imagine and add rules that do not exist.

What it looks like in the field

The difference between an overall request deadline and per-subtask deadlines also shows up in DB calls and external APIs. If four tasks of 2 seconds each run in sequence, the whole user request can far exceed 2 seconds. You have to pass the remaining budget downward and decide what to close on cancellation. This course implements only up to the budget of a single synchronous TCP frame. Large-scale concurrent connections, event loops, and backpressure on a global send queue are in the scope of a separate real-time course.

For example, if the start time is 10.0 seconds and the budget is 1 second, the deadline is 11.0. If the first byte arrives at 10.4, the budget for the next recv is 0.6 seconds, and if the next one arrives at 10.8, it is 0.2 seconds. If the byte after that takes 0.4 seconds to arrive, it was not read within the remaining 0.2 seconds. An implementation that gives a fresh 1 second every time makes every call look like a success, but it has already broken the contract for the whole request.

Conversely, if you repeatedly call the real time.sleep in a test where the caller's clock is injected, fake time and real time get mixed. The clock function only returns the current time, and the fake socket advances the fake time in step with the bytes. On the path that uses a real socket, keep the default monotonic. If you write this distinction into the API, other implementations can be tested against the same contract, and there is no need to force the test to fit a particular import statement.

After a timeout, you have to decide whether to read the rest of the frame again or close the connection. This course discards the connection after a read failure, so it does not stitch an intermediate state together through another call. If you catch the exception and return empty bytes instead, incomplete data looks like a valid empty message. Passing the kind of error to the caller is the starting point for separating the recovery policy.

What you will do in the next check

In the quiz right after this, you distinguish the limit from the overall deadline, and in the last module you run a normal split input, a mid-frame EOF, an oversized length header, and slow input, one at a time. The test also includes the plausible wrong answer that grants a fresh timeout every time. You check not only that an error occurred but also that the original socket settings came back. One second is a teaching contract and not a recommended value for every job.