Their side is down — telling six faces apart to place the blame
Goal
You build a judge and a classification table that separate connection refused, name resolution failure, connect timeout, read timeout, slow response, and partial response. You measure directly that an attempt never ends when there is no timeout, build a retry that keeps its deadline, and check six targets at once.
Why it matters
"The integration doesn't work" is at least six different incidents. This is also why it is an outage on our side while the other side's dashboard is green — a connection that could not take our hand never reaches the other side's application, so there is no record in their logs. The reason separating them matters is that the next action differs. In particular, whether you may retry splits. A failure before the connection means the other side never received the request, so you may send it again, but with a read timeout and a partial response the request may already have been processed. If that request is a transfer, a retry means a double charge. If you give no timeout at all, it is worse. The request waits until the other side answers, and meanwhile the worker thread does not come back. This is the most common path by which the other side's slowness spreads into our outage, and if retries are layered on top, the waiting multiplies. The grader does not trust your explanations. It sets up six faces on its own ports, actually hooks your judge up to them, and compares the classification. The ports change on every run, so you cannot memorize values and plug them in.
Steps
- Create and run /root/dep/gen_dep.py to create /root/dep/servers.py.
- Create /root/dep/classify.py so that it separates ok, connection_refused, and dns_error.
- Give the connect and read timeouts separately to tell connect_timeout and read_timeout apart.
- Separate the partial response from the slow response, and measure the no-timeout case in /root/dep/timeout.json.
- In /root/dep/classes.json, write the location of responsibility and the retry safety of the six faces.
- Build a deadline-keeping retry with /root/dep/retry.py and write it in /root/dep/retry.json.
- Check the six targets at once and leave them as a table in /root/dep/verdict.json.
- Report in /root/dep/summary.json and /root/dep/dep_report.md in four sections.
Notes
- Server contract:
python3 /root/dep/servers.py --role <ok|slow|hang|partial|backlog> --port P [--delay-ms N] [--ready-file F]runs only according to that role. When the ready file appears, it is ready, and to finish it you kill the process. - Judge contract:
python3 /root/dep/classify.py --url <주소> [--connect-timeout S] [--read-timeout S] [--slow-ms N] [--out <json>](where the placeholders are the address and the output path) outputs one JSON object containing url, class, elapsed_ms, evidence, connect_timeout, and read_timeout. - The six class names are
dns_error·connection_refused·connect_timeout·read_timeout·partial_response·slow_response, and normal isok. - The criterion for a slow response is
--slow-ms. If it succeeded but exceeded that time, record it asslow_response. If you record everything that returned 200 as success, you cannot see the trend of the other side getting slower. - Timeouts: requests accepts the two separately, as in
timeout=(연결, 읽기)(connect, read). If you give a single value, the two incidents get smeared into one face. If you give--connect-timeout 0or less, make it wait with no timeout. - Retry contract:
python3 /root/dep/retry.py --url <주소> --attempts N --deadline-s S [--connect-timeout S] [--read-timeout S] [--out <json>](with the address and output path filled in) outputs attempts_allowed, attempts_made, deadline_s, elapsed_ms, final_class, tries, and gave_up. When the deadline is exceeded, it must stop even if attempts remain. - Write the location of responsibility as one of three values:
우리·상대·경로(in order: ours, theirs, path). Decide retry safety by whether the request may already have been processed — a read timeout and a partial response are risky. - Reproduction materials: make a name resolution failure with the
.invalidname reserved by RFC 2606, a refused connection with a port nobody is listening on, and a connect timeout with a socket that fills the queue and does not accept (the backlog role). - Common mistakes: giving the timeout as a single value, recording everything that returns 200 as success, recording a partial response as a parsing error, and retrying a non-idempotent request after a read timeout.
- Do not build a load test. The budget for one grading is 60 seconds. Always kill the servers when you are done with them.
Set up the six faces of the integration counterpart
Create and run /root/dep/gen_dep.py to create /root/dep/servers.py. The five roles ok, slow, hang, partial, and backlog must each behave only in their own way.
Even without the internet, all six faces can be reproduced with local sockets. Save this script as it is and run it, then start one ok role and call it once with curl. Do not forget to kill the processes when you are done.
Build the judge that tells the faces apart
Create /root/dep/classify.py so that it classifies a normal response as ok, a port nobody is listening on as connection_refused, and a .invalid name as dns_error. The output must contain class, elapsed_ms, and evidence.
requests reports a failure by the kind of exception. However, a refused connection and a name resolution failure both arrive as ConnectionError, so you must look at the message to tell them apart. In evidence, leave the tail of the message that was the basis of the judgment as it is — that is what you will attach to the report later.
Separate connect from read
Output connect_timeout and read_timeout separately. The backlog role never finishes taking hands, so it is connect_timeout, and the hang role takes the hand but gives no answer, so it is read_timeout. The two timeouts must be given separately.
The judge from step 2 lumped the two timeouts into one face, timeout. requests throws a failure in the connect stage as ConnectTimeout and a failure in the read stage as ReadTimeout separately, so if you catch the two exceptions separately, they split. If you leave them as one face, you cannot tell a firewall problem from the other side's delay.
Partial response, slow response, and no timeout
Classify the partial role as partial_response, and the slow role as slow_response when it exceeds --slow-ms. And in /root/dep/timeout.json, write with_timeout, without_timeout (the exit code when it was cut off from outside after 5 seconds), and note.
A partial response shows up only when you receive the body to the end — if you look only at the headers and record success, you miss it. A request with no timeout never ends by itself, so you have to cut it off from outside. The fact that the exit code of timeout 5 python3 ... is 124 is itself the evidence.
A table of responsibility and retry for the six faces
In /root/dep/classes.json, for each of the six faces (dns_error · connection_refused · connect_timeout · read_timeout · partial_response · slow_response), write whose (우리 / 상대 / 경로, meaning ours / theirs / path), retry_safe (true/false), and next_step (at least 10 characters). The two where retry is risky are read_timeout and partial_response.
Decide retry safety by 'may the request already have been processed?' A failure before the connection was made means the other side never received the request, so it is safe, and a failure after taking the hand may already have been processed, so it is risky. In next_step, write the first place to look when you meet that face.
Build a retry that keeps its deadline
Create /root/dep/retry.py so that it honors --attempts and --deadline-s together, and leave the result of running it against the hang role in /root/dep/retry.json. When the deadline is exceeded, it must stop even if attempts remain.
A retry with only an attempt count cannot promise how long it will wait in the worst case. Before each attempt, look at the remaining time, and stop if there is none. Each attempt must also have a timeout for the deadline to mean anything — an attempt that never ends cannot be saved even by a deadline.
Check the six targets at once
Set up all six faces, judge each once, and in /root/dep/verdict.json write targets (name, role, class, whose, retry_safe) and the counts of ours, theirs, and path. All six classes must appear once each.
A checklist is not written by hand one line at a time but filled in by running the judge. Take whose and retry_safe from the table in step 5 — if the table and the checklist disagree, one of the two is out of date.
Report with the location of responsibility separated
In /root/dep/summary.json, write classes_covered, retry_elapsed_ms, retry_deadline_s, retry_attempts, theirs, path, and no_timeout_exit_code, and in /root/dep/dep_report.md, report in four sections: ## 무엇이 안 됐나 ## 우리인가 그들인가 ## 재시도는 어떻게 했나 ## 남은 위험 (in order, these mean: what failed, ours or theirs, how the retry went, and the remaining risk).
If you build this table before calling the other side, the call gets shorter. If you write a connect timeout as 'the other side's outage,' they can find nothing in their own logs — leave that distinction in the report. Also write the retry's deadline and the time it actually took as numbers.