Six meanings of "their side is down"
One-line summary
"The integration doesn't work" is at least six different incidents, and the moment you separate those six, the location of responsibility and the next action are decided together.
Why this is needed
In the field, when you hear "the other side is down" and call them, they answer that their dashboard is green. Neither is lying. What we saw on our side and what the other side sees are simply different incidents.
Failing to find the name can happen even when the other side is perfectly fine. A refused connection means the other host is alive but nobody is on that port. A stuck connect means the other side cannot take our hand, and in this case nothing is recorded in the other side's application log. A stuck read means the other side took our hand but cannot give an answer, and here the request may already have arrived on their side and be in progress. That last sentence is decisive — whether you may retry is decided here.
How it works
Laid out as a table, the six faces look like this.
얼굴 무슨 일이 일어났나 책임의 위치 재시도
dns_error 이름을 주소로 못 바꿨다 경로 안전
connection_refused 호스트는 답하는데 포트가 닫혔다 상대 안전
connect_timeout 손잡기(TCP)가 끝나지 않았다 경로 안전
read_timeout 손은 잡았고 답이 안 온다 상대 위험
partial_response 본문이 약속한 길이보다 짧다 상대 위험
slow_response 답은 왔는데 늦었다 상대 안전
The retry column is the core of this table. A failure that happened before the connection was made means the other side never received the request, so sending it again does not make the same thing happen twice. But with a read timeout and a partial response, the request may already have been processed. If that request was a payment or a transfer, a retry means a double charge. In this case you must not retry without protection such as an idempotency key.
Fortunately, the way to tell them apart is not hard. Python requests lets you give different timeouts for connect and read, and if you give them as a tuple, the failures also arrive split as ConnectTimeout and ReadTimeout. If you give them as a single value, the two incidents get smeared into one face.
If you give no timeout at all, something worse happens. With no default, the request waits until the other side answers or until the kernel gives up on the connection. Meanwhile that worker thread does not come back, and if requests pile up, we die first. This is the most common path by which the other side's slowness spreads into our outage.
If retries are layered on top, things get twice as bad. Making three attempts that never finish means waiting three times as long. So retries must come with an overall deadline. With a retry that has only an attempt count, nobody knows how long it will wait in the worst case.
What you see in the field
First, believing you cannot reproduce it without the internet. In fact all six faces can be reproduced with local sockets alone. A refused connection just needs a port nobody is listening on, and a connect timeout can be made with a socket that fills the queue and does not accept (listen(2)'s backlog is that spot). For a name resolution failure, use a .invalid name, which RFC 2606 reserved for this very purpose.
Second, mixing up the other side's fault and our fault. If you report a connect timeout as "the other side's outage," they can find nothing in their own logs. A connection that could not take our hand never even reached the other side's application. If this distinction is in the report, the investigation starts from the firewall and the path.
Third, recording slow responses only as successes. If you write down that a 200 came so it was a success, nobody sees the other side getting slower and slower. You need to put a threshold on successes too and record anything over it as a different face, so the trend becomes visible.
Fourth, recording a partial response as a parsing error. The JSON parser blew up, so you think the response format changed. In reality the body was cut off midway, which means the other side or an intermediate device dropped the connection. Comparing Content-Length with the number of bytes actually received separates them immediately.
What really matters in practice
- Give separate timeouts for connect and read. If you give a single value, the two incidents get smeared together.
- Whether you can retry is decided by the face. With a read timeout and a partial response, it may already have been processed.
- Put an overall deadline on retries. A retry with only a count cannot promise the worst case.
- Write down our fault, the path's fault, and the other side's fault separately. That line decides where the investigation starts.
What you will do in the next lab
You set up the six faces of an integration counterpart on local sockets, and build a judge that throws one request and tells those faces apart. You give the connect and read timeouts separately to see the two incidents split, measure directly that an attempt never ends when there is no timeout, and build a retry that keeps its deadline. Finally, you check six targets at once and report on one page a table that writes down the location of responsibility separately.