Connection Refused and Timeout Are Completely Different News
Summary
In a report that "the API is not working," the thing that carries the most information is not the log but the exact wording of the failure message.
Why this was needed
When a customer says "the API is dead," five entirely different situations can hide behind that sentence. And which one it is is mostly decided by a single line of the failure message.
Connection refused — this is not bad news. It is evidence that the packet went to the destination and came back. It has already passed every gate along the path, such as routing, firewalls, and NAT, and the remaining cause narrows down to inside one target host. The process is dead, or it is listening on a different port, or it is bound not to 0.0.0.0 but only to 127.0.0.1.
Connection timed out — it carries no information. It means nobody answered, and a firewall may have silently dropped it, there may be no route, or the other side may be overloaded. The firewalls used in practice almost always have a silent-drop policy, so when a firewall blocks, you get a timeout, not refused. So looking through the firewall after seeing refused means re-checking a gate that has already been passed.
Name or service not known — not a single packet went out toward the target server. Nothing in the firewall log is normal, and the investigation goes toward name resolution.
If you also measure the time it took to fail, you gain confidence. If it ended in a few milliseconds, the round trip completed (refused), and if it was cut at about 2 minutes 7 seconds, the kernel has exhausted all its SYN retransmissions. If you set the app timeout to 30 seconds but it actually failed at 127 seconds, that is a sign that setting is not being applied.
How it works
Once connected, the next thing is the status code. The ones most often misdiagnosed here are 502 and 504.
A 502 is not a case of not receiving a response but of receiving an invalid response. A case where a response did arrive but the proxy could not parse it is also a 502. So increasing the timeout for a 502 has no effect at all. A case where no response was received in time is a 504.
And when attaching to an API with poor documentation, there are three more things you must check.
How do you know where pagination ends? If there is an explicit flag such as has_next, you trust it, and if not, you loop until an empty array comes back. A common bug here is to compute the number of pages from only the first page's total and then actually miss the last page. It is the classic mistake of forgetting the remainder of a division.
How do you handle transient failures? A 503 or network error is often moved on to retry. But what may and may not be retried separates. A lookup is the same however many times you repeat it, but retrying a create request produces duplicates. A 4xx gives the same result however many times you send it unless you fix the request, so it is not a retry target.
How many gaps are there that are not in the documentation? This is what most often catches people in practice. It is common for a field marked required in the specification to arrive as an empty string in the actual response, and if you put that value into the aggregation as is, a silently wrong number comes out.
What it looks like in the field
So the first job when attaching to a new API is not writing code but a full survey. You go through every page once and count the total count, the sum, and the number of missing values in each field.
This single survey removes weeks of later debugging. If in the first week you ask "of the 60 records in total, 6 have an empty region. How should we handle them?", later no report will come that the regional totals do not match.
Making code that leans on someone else's API safe
Once the full survey tells you what you are dealing with, the next job is to make it so we do not collapse even when that API wobbles. Since we cannot fix someone else's API, writing the assumptions into the code is the only defense.
Do not trust what you receive as is. We have already seen fields marked required in the specification arrive empty. So check the type and range at the point of reading, and if they do not match, skip that record but count and record the number skipped. If you drop them silently, later you cannot find why the total does not match, and if you fail as a whole, one record stops everything.
Always set a timeout. Many libraries have no default or an infinite one. If a single call with no timeout holds a worker forever, the connection pool exhaustion we saw earlier reproduces as is. If you can set the time to connect and the time to respond separately, set them separately.
Retry only what is idempotent. As said before, a lookup is the same however many times, but a create is not. And if you do not mix randomness into the retry interval, when the other side wobbles and then recovers, every client piles in at the same moment and knocks it down again.
Save the response as is. If you keep the original, you can rerun it when you later fix the parsing rules. If you save only the parsed result, by the time you discover the rule was wrong, the original is already gone. The cost of fetching it again is usually far larger than the cost of storing it.
Finally, put in a device to notice when the other side has changed. If you make it alert when the response's field set or record count is very different from usual, then when an API changes silently you know before your users do. Someone else's API changes without notice, and even when it gives notice, that email usually does not come to us.
What you will do in the next lab
You attach to an order API whose documentation is only a one-page wiki, check the version, go through all the pages to get the total count and sum, count the gaps not in the documentation, and get through an unstable endpoint with retries.