Who called what, and when — audit log, schema validation, timeouts
In one line
A production MCP server needs four more things: an audit log that records every call, validation of arguments against the schema before execution, a per-tool execution time limit, and call-rate limits for destructive tools. The specification makes the first two and the last a MUST for the server, and puts timeouts as a SHOULD on the client; in practice it is safer for the server to have all four.
Why this was needed
After an incident, the first question is "who called what, when, and with which arguments?" Without an answer, you fill the cause with guesses, and a server fixed by guessing causes the same incident again. The previous module's incident record could fill in cause= only because request logs happened to exist, and in real operations you must not rely on that luck.
The security considerations section of the tools specification says servers must validate all tool inputs, implement access control, rate-limit invocations and sanitize outputs (MUST). For clients it recommends (SHOULD) user confirmation for sensitive operations, showing the inputs before calling, validating results, timeouts on tool calls, and logging tool usage for audit. This module moves that list into server-side code — because the server cannot choose which host the client is.
How it works
The audit log is per call. You wrap the one function that executes a tool, measure the start time, decide success from the isError of the result, and append one line of JSON to a file. What to record: when (ts), what (tool), with which arguments (arguments), whether it worked (ok), how long it took (duration_ms), and, if it failed, why (error). Calls rejected with a protocol error are recorded too — a client that keeps sending bad arguments is a signal in itself. If you take the path from an environment variable, a grader or a test can run against a temporary file.
Schema validation comes before execution. The inputSchema you publish through tools/list is a promise made to the model, and validation is the server keeping that promise itself. If required is missing or a type does not match, you do not execute and answer with -32602 Invalid params from JSON-RPC 2.0. The tool did not fail while running; the request was wrong, so it is a protocol error, not isError. One thing to watch in Python: bool is a subtype of int, so isinstance(True, int) is true. A true arriving where an integer belongs has to be blocked separately.
TYPES = {"string": str, "integer": int, "boolean": bool}
for key in schema.get("required", []):
if key not in args:
raise RpcError(-32602, f"Invalid params: missing {key}")
for key, spec in schema["properties"].items():
if key in args and not isinstance(args[key], TYPES[spec["type"]]):
raise RpcError(-32602, f"Invalid params: {key} must be {spec['type']}")
The server sets the timeout itself. The timeouts section of the lifecycle says to set a timeout on every request sent, to prevent hung connections and resource exhaustion, and to enforce a maximum timeout even when progress notifications arrive. That is about the client side, but the server also puts a limit on its own tools for the same reason — because one stuck tool holds up a whole single-threaded server. If you use the standard library's signal.alarm to receive SIGALRM after N seconds and raise an exception in the handler, you get out even in the middle of time.sleep or a slow query. When the call ends, release it with signal.alarm(0). A call that exceeds the time is answered with isError: true and recorded as a failure in the audit log.
Also send logs through the protocol. The logging utility defines how a server declares the logging capability and sends structured logs with notifications/message notifications. The level is one of the eight RFC 5424 levels (debug, info, notice, warning, error, critical, alert, emergency), with a logger and data attached. Being a notification, it has no id. If stderr logs are for people, this is a log the host receives in structured form to show on screen or collect. You can send one line before the response for each call.
Rate limiting starts with destructive tools. A session is one process, so you can count calls per tool in a module-level dictionary. If you cap delete_order at 3 per session, an agent caught in a bad loop that would delete every order stops at three. A call over the limit deletes nothing and is answered with isError: true.
What it looks like in the field
Once you have an audit log, you need a tool to read it. A single script that counts calls and failures per tool can answer "how many times was delete_order called in the last hour and how many were rejected", and those numbers are the raw material for alert rules. If the failure rate suddenly rises, you changed the schema but did not fix the model-side description; if timeouts increase, the database behind has slowed down.
Timeout values differ by tool. Two seconds is enough for a lookup, but report generation may need 30 seconds. The order is to start with a single value and then split it per tool after looking at the distribution of duration_ms in the audit log — if you guess a different value for each tool from the start, numbers with no basis are left in the code.
What you will do in the next lab
You add, one after another, an audit log, schema validation, a slow_report tool with a 2-second limit, log notifications and a 3-call limit on delete_order to the previous module's server, and at the end you write a script that reads the audit log and summarizes calls and failures per tool.