TT Lab
Get started
Learn Learning paths Courses

The agent dropped my database

Before the agent drops your database — allowlist, read-only, confirmation

Continue in TT Lab

In one line

Three layers make a destructive tool safe. Get rid of the catch-all tool and switch narrow tools on and off with an allowlist; where writing is not needed, open the connection itself as read-only; and make a tool that deletes refuse to run without confirmation. Turn the specification's sentence that "a human must be able to deny" into code, and these three are what you get.

Why this was needed

The ingredients of an incident are always the same: a single run_sql built for convenience. If you let the model write SQL, you do not need to build ten tools, and the demo works astonishingly well. Then one day "clean up last month's test orders" turns into DROP TABLE orders. The model merely picked the shortest path, and the server put no door on that path.

The "Security and Trust & Safety" section of the specification overview deals with this situation head on. Tools represent arbitrary code execution and must be treated with appropriate caution; tool descriptions and annotations should be considered untrusted unless they come from a trusted server; and hosts must obtain explicit user consent before invoking a tool. The tools specification is more concrete: there should always be a place where a human can deny a tool call (SHOULD), and servers must validate all inputs, implement access control and rate-limit calls (MUST). The specification itself notes that it cannot enforce this at the protocol layer, so if the implementer does not put it in, it exists nowhere.

How it works

First, narrow the tools. Instead of run_sql, have one tool per intent, such as count_orders, list_customers and delete_order. A narrow tool has a narrow argument schema, and a narrow schema means a small space of requests the model can construct. The only argument of delete_order is id, so there is simply no "way to drop a table". This is the first layer of access control, and it is more reliable than any other check — there is nothing to check.

Second, switch tools on with an allowlist. The server separates all the tools it defines (ALL_TOOLS) from the tools it enables now (allowlist.json). tools/list returns only what is on the list, and tools/call also answers a name outside the list as an unknown tool (-32602). If the list file is missing or broken, no tools are enabled — a closed default. In production, you use it like this: enable only the read tools, and enable the delete tool only during the working hours when it is needed. It has to be "if it is not on the list, the tool does not exist", not "pick which of the existing tools to turn off", so that a newly added tool is not opened by mistake.

Third, apply read-only at the connection. Checking whether the SQL string starts with SELECT leaves you an endless list of things to check. You would have to follow multi-statement input glued together, comments and case variations, and even read statements that start with WITH (SQLite also allows WITH ... DELETE), and the one you miss is the incident. The answer is to open the database with mode=ro in the URI supported by Python's standard-library sqlite3.

con = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
con.executescript("DROP TABLE orders;")   # sqlite3.OperationalError: attempt to write a readonly database

Because the write is rejected by the database engine, no matter how the string is dressed up it does not get through, and the server code has no list to check at all. The server only has to catch that exception and return it as isError: true. Switch it on with one environment variable (MCP_READ_ONLY=1), so that in deployments that only need queries there is no write path at all.

Fourth, get confirmation for destructive tools. delete_order deletes only when the confirm argument is the boolean true. Otherwise it deletes nothing and returns, as isError: true text, what would have disappeared had it deleted (order number, status, amount). The host shows that text to a person, and if the person approves, the model calls again with confirm: true. This round trip is what the specification calls "human in the loop". You must not accept the string "true" — if the schema says boolean, compare with is True. You can attach hints such as readOnlyHint and destructiveHint to a tool definition's annotations, but the specification says a client must not trust these annotations unless they come from a trusted server (MUST). A hint is for display, not a safety mechanism.

What it looks like in the field

Recovery after an incident comes from backup. But the incident record matters more: what was the cause (tool design), which controls were missing (allowlist, read-only, confirmation), and what will change. Without those three lines, the next server also starts with run_sql. That is why this module's lab goes in the order reproduce the incident, recover, write the record, then add the three layers of safeguards.

One more thing. After you add a confirmation step, the complaint comes: "the agent asks every time, so it is slow". The answer is not to remove the confirmation but to narrow the tool further. A tool like cancel_test_order (test orders only, status change only) rather than delete_order is safe enough that no confirmation is needed. Safety comes not from the number of confirmation prompts but from the size of what the tool can do.

What you will do in the next lab

You will actually send DROP TABLE orders to a run_sql server, watch orders disappear, recover, and write an incident record. Then you add read-only mode, an allowlist and a confirmation argument one after another, and show that the same attack is stopped in three places.