FDE Capstone: The Warehouse Got the Same Order Three Times
NO_PROXY is not a standard: three tools, three interpretations
In one line
Proxy environment variables are just a convention and not a standard, so curl, Python urllib, and requests each read the same NO_PROXY line differently. At a customer site, before writing the settings "correctly", the first thing is to measure the actual route per tool and leave it as evidence.
Why this was needed
On the first day of delivery, the customer's security team sent three lines of settings. "It's our internal standard. Put it in as is."
HTTP_PROXY=http://127.0.0.1:3128
HTTPS_PROXY=http://127.0.0.1:3128
NO_PROXY=corp.example,127.0.0.0/8
After putting it in, the integration program's health check (curl) was green, but the metrics collector written in Python started getting 403 from the internal API. A requests call that should have gone out to an external SaaS instead skipped the proxy and died at name resolution. The three programs read the same three lines and moved in three different ways. Nobody is wrong. http_proxy is just a convention shared by several tools, and the curl, Python, and requests documentation quoted below each explain only their own rules. Since there is no place that fixes a common grammar, the rules differ slightly from one implementation to another, and that difference shows up all at once inside the customer network.
It is risky for an FDE, in this situation, to say first "You just need to change the settings like this". The customer's owner will check with curl and answer "Ours is normal", and that statement is also true. What you need is evidence that shows, in a record, which request went through the proxy.
How it works
Below are the results we measured ourselves with a recording fake proxy running on this lab image (measured: curl 8.5.0, Python 3.12.3, requests 2.32.3).
1. The case of variable names. The ENVIRONMENT section of the curl manual says that it reads both lowercase and uppercase, but http_proxy alone is accepted only in lowercase. Why this exception is needed is explained by the urllib.request documentation — in a CGI environment, the Proxy: header sent by a client can be injected as an environment variable called HTTP_PROXY. That is why Python ignores the uppercase HTTP_PROXY only in a CGI environment where REQUEST_METHOD is set, and otherwise accepts uppercase too. It behaved that way in our measurement as well. As a result, if you set only an uppercase HTTP_PROXY, curl goes direct and the two Python tools go through the proxy. If both lowercase and uppercase are present, lowercase wins — the curl manual and the urllib documentation say the same thing about this.
ALL_PROXY also splits them. The curl manual says it uses this when there is no per-scheme variable, and the requests documentation also lists all_proxy as a variable it reads. In our measurement the two went through the proxy, but urllib went direct when only ALL_PROXY was set.
2. Name matching in NO_PROXY. The curl manual explains that it matches each entry as "the host itself or a name belonging to that domain", and gives the example that local.com matches www.local.com but does not match www.notlocal.com. urllib also respects the dot boundary. requests is different — in our measurement, NO_PROXY=corp.example sent even saascorp.example direct. This is because it looks only at whether the end of the string is the same. A leading dot (.corp.example) matched api.corp.example and corp.example itself in all three tools and did not match saascorp.example. On the other hand, an entry with an asterisk attached like *.corp.example matched nothing in any of the three tools. An asterisk means "everything direct" only when the whole list is the single character *, and if you mix it in like *,localhost, the asterisk loses its power (measured).
3. IPs, CIDR, and ports. The manual says that curl accepts CIDR notation such as 127.0.0.0/8 from 7.86.0. requests also applied CIDR to IP address hosts. urllib does not know CIDR — it compared only as strings and sent 127.0.0.2 through the proxy. Ports are the other way around. The urllib documentation says it allows an entry with a port attached like some.host:8080 and it actually matched, but curl could not match an entry with a port, and requests could not match when a port was attached to an IP but matched when a port was attached to a name. Also, NO_PROXY=127.0.0.1 did not match localhost. This means none of the three tools resolved names to addresses for comparison (measured).
4. A proxy passed by the code. The Proxies section of the requests documentation says that a value put into session.proxies can be overridden by the proxy in an environment variable, so to be sure it is used, pass it with the proxies argument on each request. But NO_PROXY was not applied to the proxies passed on each request (measured). A connector that passes the proxy from a config file as an argument cannot go directly to the internal API no matter how you fix NO_PROXY.
| Same condition | curl | urllib | requests |
|---|---|---|---|
Uppercase HTTP_PROXY only |
Direct | Proxy | Proxy |
NO_PROXY=corp.example → saascorp.example |
Proxy | Proxy | Direct |
NO_PROXY=127.0.0.0/8 → 127.0.0.2 |
Direct | Proxy | Direct |
NO_PROXY=127.0.0.2:8080 → 127.0.0.2:8080 |
Proxy | Direct | Proxy |
What it looks like in the field
On a customer network, the cause always shows up as a different symptom. If the proxy rejects an internal destination with 403, it looks like "API authentication is broken", and if an outside call skips the proxy, it looks like "DNS is strange". If the health check is green because it is curl, everyone suspects the application. At such a time, one line of the proxy access record ends the argument — did that request arrive at the proxy or not.
The second most common picture is "the security team's standard" colliding with "the reality of the tools". You cannot ask them to fix the standard just because a urllib-based tool does not understand the internal range written in CIDR in the standard document. The FDE keeps the standard and writes in one more set of the common denominator that all three tools understand (leading-dot domains, an explicit IP list, both lowercase and uppercase), and leaves the reason on record.
What really matters in practice
- Do not infer the settings; measure them. If you start a fake proxy that keeps records locally, you can confirm each tool's decision even without access to the customer's proxy.
- Use the common denominator: both lowercase and uppercase, a leading dot for domains, both CIDR and individual addresses for IPs, and do not attach ports.
- If the code passes the proxy directly, the code has to take care of
NO_PROXY. It also takes care to keep other proxies in the environment from getting mixed in. - Leave the check results with the tool names and versions. When a version changes, the rules have also changed before (curl's CIDR support is an example).
What you will do in the next lab
You start a recording fake proxy and an internal API, reproduce the failure of the customer's settings, and measure 18 prepared cases directly with the three tools and leave them in a table. Then you make settings under which the three tools go the same way, fix a connector that uses the proxies argument, and finally build a probe that measures the actual route per tool and judges it for any settings file.