TT Lab
Get started
Learn Learning paths Courses

The agent dropped my database

Add audit logging, schema validation, timeouts and rate limits

Continue in TT Lab

Goal

You add four things needed in production to the server you locked down in the previous lab: an audit log that records every call, validation of arguments against the schema before execution, a per-tool execution time limit, and a call-count limit for destructive tools. You also write a tool that reads that log and summarizes it.

Why it matters

When an incident happens, the first question is "who called what, when, and with which arguments?" Without an audit log you cannot answer it, and then the same incident happens again. The specification says servers must validate all tool inputs and rate-limit invocations (MUST), and that clients should put timeouts on tool calls and log usage for audit (SHOULD). This lab moves those sentences into code as they are. In particular, you must not rely on the client alone for timeouts — if the server does not cut off by itself, one stuck tool holds up the whole session.

Steps

  1. Save /root/mcp/audit/seed.sql and load it into /root/mcp/audit/shop.db. There must be 5 rows in customers and 8 rows in orders.
  2. Create /root/mcp/audit/server.py. It answers initialize and offers three tools in tools/list: list_customers, count_orders and delete_order (confirm required). The DB path is MCP_DB (default /root/mcp/audit/shop.db).
  3. Record every tools/call as one line of JSON in /root/mcp/audit/audit.jsonl (changeable with the environment variable MCP_AUDIT_LOG). The keys are ts, tool, arguments, ok and duration_ms, and if the tool returned isError, ok is false.
  4. Check the arguments against inputSchema before running the tool. If a required field is missing or a type does not match (for example, a number for status or a string for id), do not run it and answer with the JSON-RPC error -32602.
  5. Add a tool slow_report (argument seconds, integer) and limit the execution time of a single tool to 2 seconds (environment variable MCP_TOOL_TIMEOUT, default 2). If it is exceeded, answer with isError: true and text containing timeout, and record it in the audit log with ok: false. seconds: 1 must end normally.
  6. Declare logging in the capabilities of initialize, and for every tool call send a notifications/message notification (level is one of the RFC 5424 levels, plus logger and data) to stdout before the response. A failed call is at the error level.
  7. Allow delete_order only 3 times per session. The fourth call deletes nothing and answers with isError: true and text containing rate limit.
  8. Create /root/mcp/audit/summary.py. It reads audit.jsonl and writes {"calls": n, "failed": m} per tool to /root/mcp/audit/summary.json (to that path if the environment variable MCP_SUMMARY_OUT is set). The real audit.jsonl must have accumulated at least five lines of call records.

Notes

Create the shop DB

Save /root/mcp/audit/seed.sql and load it into /root/mcp/audit/shop.db. There must be 5 rows in customers and 8 rows in orders.

With sqlite3, sqlite3 shop.db < seed.sql runs the whole file. In Python, use sqlite3.connect(...).executescript(open(...).read()). Loading again into a DB that already exists gives an error saying the table exists, so delete it first.

Set up the base server

Create /root/mcp/audit/server.py. It answers initialize and offers three tools in tools/list: list_customers, count_orders and delete_order (confirm required). The DB path is MCP_DB (default /root/mcp/audit/shop.db).

Take v3 from the previous lab and just remove the allowlist (this lab enables all three tools). delete_order does not delete unless confirm is true.

Record every call

Record every tools/call as one line of JSON in /root/mcp/audit/audit.jsonl (changeable with the environment variable MCP_AUDIT_LOG). The keys are ts, tool, arguments, ok and duration_ms, and if the tool returned isError, ok is false.

Wrap the one function that executes a tool — measure the start time, decide ok from the result's isError, and append one line. The shape of one line is {"ts": "2026-01-01T00:00:00+00:00", "tool": "get_weather", "arguments": {"location": "Seoul"}, "ok": true, "duration_ms": 12, "error": null}. The grader gives MCP_AUDIT_LOG a temporary path and, after two calls (one normal, one with an unknown status value), checks that there are two lines.

Check the arguments before running

Check the arguments against inputSchema before running the tool. If a required field is missing or a type does not match (for example, a number for status or a string for id), do not run it and answer with the JSON-RPC error -32602.

A small function is enough that maps the schema's required list and the type in properties to Python types (string→str, integer→int, boolean→bool). A check failure happens before the tool runs, so it is a protocol error (-32602, Invalid params).

Put a limit on tool execution time

Add a tool slow_report (argument seconds, integer) and limit the execution time of a single tool to 2 seconds (environment variable MCP_TOOL_TIMEOUT, default 2). If it is exceeded, answer with isError: true and text containing timeout, and record it in the audit log with ok: false. seconds: 1 must end normally.

If you set signal.alarm(TOOL_TIMEOUT) and raise an exception in the SIGALRM handler, you get out even in the middle of time.sleep. When it ends, signal.alarm(0). The grader sends seconds: 6 and measures whether an isError response arrives within 5 seconds.

Send logs through the protocol

Declare logging in the capabilities of initialize, and for every tool call send a notifications/message notification (level is one of the RFC 5424 levels, plus logger and data) to stdout before the response. A failed call is at the error level.

A notification looks like {"jsonrpc":"2.0","method":"notifications/message","params":{"level":"info","logger":"...","data":{...}}} and has no id. Just write one line before you write the response. Unlike a stderr log, this is a log the client receives in structured form.

Limit the number of calls to a deleting tool

Allow delete_order only 3 times per session. The fourth call deletes nothing and answers with isError: true and text containing rate limit.

A session is one process, so you can count calls per tool in a module-level dictionary. The grader calls it four times with confirm=true against a temporary copy of the DB and checks that exactly 3 rows disappeared.

Read the audit log and summarize it

Create /root/mcp/audit/summary.py. It reads audit.jsonl and writes {"calls": n, "failed": m} per tool to /root/mcp/audit/summary.json (to that path if the environment variable MCP_SUMMARY_OUT is set). The real audit.jsonl must have accumulated at least five lines of call records.

Run json.loads on each line, count by tool, and add to failed when ok is false. If you accept the log path as the first argument, the grader can check against a temporary log. If you ran the earlier steps, audit.jsonl already has several lines.