TT Lab
Get started
Learn Learning paths Courses

The snack machine died before ACK

Reject a Late Courier's Completion Stamp

Continue in TT Lab

In one line

A lease is a right to hold something briefly, and the claim token distinguishes which round of the right it is. A stale worker must not be able to change the new worker's records even if it comes back alive.

Why this was needed

Dispatcher A took a snack order and then stalled. If you completely deleted the order from the queue and kept it only in A's memory, the order would vanish along with it. Conversely, if you leave it visible in the queue as it is, B can take it at the same time. That is why you store a leased state and put a marker saying it was entrusted to A for a fixed time. If it is not completed after the time passes, another worker can take it again.

The problem is that expiry does not terminate A. If A was only waiting a long time for a network response, then after B newly claims it, A can also receive a success response. If A then marks it complete with only the id as the condition, it hijacks B's current work. Checking only the name of whoever is running is not enough either. This is because when a worker with the same name restarts, you cannot tell the earlier run from the new run.

How it works

A jobs row has owner, lease_until, token, and attempts. When a claim succeeds, attempts and token each go up by 1. owner is a worker name that can be explained, and token is a monotonically increasing number that distinguishes the claim generation of that work. One claim result contains id, qty, owner, token, attempt, and lease_until. finish reflects the result only when the marker it was given matches the current row entirely and the lease and the work's deadline still remain.

0ms     A가 선점: token=1, lease_until=1000
1000ms  B가 재선점: token=2
1001ms  A가 token=1로 완료 요청 → False, 현재 상태 유지

The boundary is valid only when now < lease_until. If now equals the expiry time, the right has already ended. A new claim, conversely, can take work with lease_until <= now. If you interpret the two conditions differently, two parties may believe they are valid at the same instant, or a gap opens in which no one can take it. The test feeds in before expiry, exactly at the expiry time, and after expiry.

The claim ties the query and the update into one write transaction. If you read the available work first and start the transaction later, A and B can read the same pending row. This queue serializes the selection and the token update with SQLite's short BEGIN IMMEDIATE transaction. It is a work queue with no reason to hold a read-point snapshot for long, so it uses a rollback journal. The concurrency shape it needs differs from the WAL snapshot export of the earlier lab.

After committing the claim, it releases the lock and sends. If you wait for the network response inside a DB transaction, new intake unrelated to this work and claims by other workers are blocked too. You have a separate connection actually accept a different work item inside the test's send callback to check whether the lock was left behind. What gets graded is whether the intent of a short transaction that does no external work is kept in actual execution.

What it looks like in the field

SQS's visibility timeout is also a related concept that hides a received message from other consumers for a fixed period. However, that product's delivery guarantees and this local SQLite queue's contract are not the same. This lab neither implements nor calls SQS. The official explanation also says the visibility timeout is not a guarantee that completely eliminates the possibility of duplicate delivery, so you must think separately about the idempotency of the receiving work.

Here the claim token rejects stale completion markers inside the queue. It does not implement fencing of an external resource where the receiving server checks the token. If A's request has already reached the receiving server, then even if the queue rejects A's finish, that external effect is not undone. That is why, in the comprehensive test, the receiving server records the business ID and quantity together and does not add the effect again for a resend of the same ID. It is the reason the queue's ownership and the receiving effect each need their own boundary.

If you give a very long lease, the time before you can hand a dead worker's job to someone else grows, and if you give a very short one, you needlessly send again the work of a worker that is still healthy. This implementation does not build a heartbeat that renews the lease. When you run long tasks, you have to design the processing time, the renewal interval, safe cancellation, and receiver idempotency together. Simply changing lease_ms to a bigger value is not a complete solution.

You also keep the lease from becoming longer than the original work deadline. For example, if deadline=100 but now=0 and lease_ms=1000, the actual lease_until is 100. The lease stored in the queue must not contradict the work's budget. Expired work whose attempts are exhausted is quarantined before it is given a new token, and work that already holds a valid lease is not handed to another worker by repeated queries alone.

What you will do in the next check

In the quiz that follows, you judge what each of the name, the business ID, and the claim number distinguishes. In the comprehensive lab after it, you start two real child processes at the same time and look to see that there is only one claim result. The failure message at the completion boundary shows the current result and the expected result together, and you check that the old token is rejected even when it is re-claimed under the same worker name.

Reference: SQLite transactions, SQS visibility timeout. A numeric token is not an access permission or an authentication mechanism.