TT Lab
Get started
Learn Learning paths Courses

Air-Gapped Sites — Defence and Government

If the clocks are wrong, the records are not evidence

Continue in TT Lab

In one line

An air-gapped network has no outside time source. A time source set up inside the network aligns only the relative order among hosts and does not guarantee absolute time, and a statement that says "the records are accurate" without knowing that difference falls apart in an audit.

Why this was needed

Both outage investigations and audit responses ultimately ask "what happened first?" But the time on a record is the value shown by the clock of the machine that wrote it, and every machine's clock is different. If they differ on the scale of seconds a person notices, but a difference of about 2 seconds flips cause and effect without anyone noticing. In real investigations this is the most common place where things go wrong. If a deployment looks as if it happened after the outage, the investigation goes off in the wrong direction.

Where the internet reaches, this problem is usually solved quietly. If you watch several public NTP servers, clocks converge on their own. In an air-gapped network that premise disappears. You set up one time source inside the network and everyone aligns to it, but that time source itself has nothing outside to align to. In the stratum defined by RFC 5905, such a time source stands in the position of taking itself as the reference. As a result, all the records in the network are consistent with one another, but nobody can say how far ahead or behind the whole set is relative to the outside world. If you do not write this fact in the statement, it becomes a problem later.

How it works

The judgment is built in four layers.

First, gather the notations into one. Even within the same system, time notations in logs vary. ISO 8601 with an offset, local time without an offset, epoch seconds, and UTC with a Z at the end. The most dangerous one here is a time without an offset. If you read it as UTC, in the Korean time zone it is shifted by a full nine hours. You cannot tell which it is from the log file alone, so that fact has to be taken from the system documentation, and the statement must also state that basis. This is the reason RFC 3339 makes the offset mandatory.

Second, follow the hierarchy. If you follow step by step where each host gets its time from, most reach the declared time source. The hosts that do not reach it are the problem. A host whose configuration was erased and which is using its own clock as is keeps stamping seemingly perfectly good times, so it cannot be distinguished from the log alone. Following the hierarchy to the end is the only way.

Third, find the places where records collide. In requests and responses exchanged between two hosts, if there is a case where the response comes before the request, it means the two clocks are apart by at least that difference. This is not a guess but a lower bound. It is a value that arises because the sending side and the receiving side stamped the same event with different clocks, and comparing it with the offsets in the synchronization report also confirms whether the report is true.

Fourth, look at what remains after correction. If you correct by the offsets, most contradictions disappear. If something remains that does not disappear, it is not a clock problem but a defect in the record itself. Telling these two apart is the core of this work. Stretches explained by correction can be used conditionally, and stretches that are not explained cannot be used as evidence.

One more thing. Clock synchronization does not prove that that record existed at that time. That is the work of another layer, and it needs a structure where a third party signs, like the timestamp of RFC 3161. A statement that speaks of synchronization and timestamps as the same thing loses trust. On log management in general, NIST SP 800-92 treats time synchronization as a premise of log reliability.

What it looks like in the field

At one air-gapped delivery, the investigation spun its wheels for two days. It was because a batch job looked as if it had been processed before the time it was put into the queue. The cause was that the processing node's clock was 2.4 seconds fast, and the synchronization report had that value written in plain sight. Nobody had just put that report side by side with the logs. After correction, exactly one contradiction remained, and that one was the real bug.

Another thing I often see is a host that lists itself as its own time source. Its configuration was erased once during installation and never restored, and that host's logs keep stamping plausible times. It does not show until you draw the hierarchy as a picture.

The third is the wording of the statement. The temptation to write "the times are accurate" when handing over investigation results is great, but the data does not back that sentence. What the data backs is only up to "the records in this stretch do not contradict one another and are explained by the reported offsets." This difference is not wordplay. If even one record later goes off, a statement that wrote the former sentence loses trust as a whole, while a statement that wrote the latter only needs that one record looked at again. The documents auditors like are not documents that assert but documents in which what was confirmed and what could not be confirmed are separated.

What you will do in the next lab

You gather the logs of five hosts, split across four notations, onto one axis, and count how many events shift and by how much if you wrongly read the logs without offsets as UTC. Next you follow the time hierarchy to find hosts that do not reach the declared time source, find the lower bound of the clock skew from the contradictions between requests and responses, and then correct. At the end you grade how far each host's records can be trusted, and write a statement that states together what can be trusted and what cannot.