Production Backend API Capstone
Metrics, Logs and Traces Answer Different Questions
One-line summary
Metrics explain overall trends and alerts, structured logs explain the context of a single event, and trace identifiers explain causality across boundaries. The three signals must be linked through safe correlation fields that come from the same request.
Why a single log line cannot resolve an outage
A string like error happened does not tell you how often, on which path, or in what state something failed. Metrics for the number of orders created and HTTP latency show changes and thresholds. The event, route, status, and duration_ms in a JSON log let you search for individual requests. If the same trace ID is in the response header and the log, you can find exactly the request a user reported. If there are several services, propagate this ID to the next call.
Metric names should reveal their unit and cumulative meaning, like orders_created_total and http_request_duration_seconds. Using a user ID or order ID as a metric label causes a cardinality explosion. Put such unique values in logs if needed, subject to personal-data and secret-handling policies. Never record auth headers, cookies, passwords, connection strings, or internal service DNS in any signal.
How to verify in practice
The grader sends a create request to the server to get the trace ID from the response, and reads the two key metrics from /metrics. After the server exits, it parses each log line as JSON and checks that a single record contains the same trace ID, the /orders path, status 201, and a numeric duration_ms. Merely having the strings json or trace_id in the source is not evidence. Output at import time, or raw token text appearing in the logs, also fails the check.
Dividing up where each of the three signals is used
"Collect logs, metrics, and traces" is a common refrain, but unless you decide which one to look at for which question, you will flounder during an outage even with all three.
| Question | What to look at | Why |
|---|---|---|
| Is there a problem now | Metrics | Values are cheap and cover everything. Set alerts only here |
| Where is it slow | Traces | Time spent in each segment a single request passed through |
| Why did it happen | Logs | Context at that moment. The most expensive, so look last |
Set alerts only on metrics. If you alert on a log string, changing the message slightly makes the alert silently stop firing, and you find that out during an outage.
Link the three signals with one identifier. Create a trace ID for every request, put it in the log, and attach it to the metric as an exemplar as well. Then you can go from "one slow request" straight to that request's logs. Without this link you end up searching logs by timestamp, and that approach becomes useless the moment traffic grows.
Cardinality is a problem only for metrics. Putting a user ID in a log is fine, but putting it in a metric label makes the time series explode. Keep high-cardinality values in logs and traces, and only low cardinality in metrics.
Measure once more from the user's side. Even when the server returns 200, the user can still fail. If you have a value measured in the frontend or a synthetic monitor running from outside, the situation where "internal metrics are normal but complaints keep coming in" goes away.
The criterion for deciding what to keep is "what decision will this drive?" A metric that changes no decision clutters the dashboard and only costs money. Each time you build a dashboard, write one line about the action you would take on seeing that screen, and half of them disappear.
Practical judgment criteria
Good observability is not printing many after-the-fact debug statements; it is deciding the operational questions first and then designing low-cost signals. Watch the success rate, error rate, latency, and traffic as metrics, and narrow down specific anomalies with traces and logs. In the deployment lesson that follows, you design probes and runtime security boundaries so that traffic is not sent to an unready Pod even when these signals look normal.