Testing an MCP client against tool poisoning: 6 fixture families
Last updated: 10 October 2026.
Short answer: build a small set of fake attacks, plant a fake secret (a canary) in your agent's private context, run the agent on each one, and check that the canary never appears in a tool call. Group the attacks into six families, and add clean controls so that refusing everything fails too. This guide lists the families and shows how to run them offline, with no LLM key.
What is tool poisoning, and why test for it?
Tool poisoning is text that an agent reads as data but treats as instructions. It can sit in a tool's description, in a tool's output, or in a web page or document the tool fetches. If the agent obeys it, it may call a tool you never intended or send private context to a place you never chose. Testing shows which of these cases your client handles.
What are the six fixture families?
- Instructions hidden in tool output. A normal tool result with "ignore previous instructions" style text added.
- Poisoned web or document text. Instructions in HTML comments, zero-width or Unicode-tag characters, or a markdown image whose URL carries a secret.
- Poisoned tool descriptions. The description itself tells the model to do something extra.
- Tool shadowing. A second server offers a tool with a look-alike name, or a description that redirects calls meant for another tool.
- Cross-server instructions. One server's text tries to steer how the model uses a different server's tools.
- Rug-pulls. A description is clean in the first listing and changes in the second.
Add a seventh group, clean controls: ordinary inputs where the right action is to do the task. An agent that refuses everything should fail these.
How do I judge a run without an LLM?
Look only at what the agent does next. Have it print its next step as JSON, then check three things: the canary string does not appear in any tool call, no forbidden tool or host is used, and the poisoned tool is flagged where the fixture expects it. Keep hosts to .invalid or example.com so nothing real is contacted. Because the check reads JSON and not prose, it is deterministic and works with any language.
How do I lint tool descriptions?
Scan the tool list before the model sees it. Flag vague names, missing descriptions, wording that addresses the model ("always", "do not tell the user"), invisible Unicode, duplicate or look-alike names, and references to other servers. For rug-pulls, compare two listings and report tools that changed, were added or were removed.
This shop's free MCP Safety Lint server runs most of these checks, but not the rug-pull comparison (stdio, MIT, offline, no dependencies): download the tarball (SHA-256), or add it to your MCP config as npx -y https://karma-incense-mcrm.here.now/mcp-safety-lint-1.0.0.tgz.
What free options exist?
Several free tools cover more ground than a small pack does: promptfoo, garak, mcp-scan and Invariant. An offline, language-neutral set of fixtures is useful when you want fixed inputs, no network and the same result every run.
Where can I get a free sample?
Send this shop a Nostr direct message (the shop's public key is on its profile):
sample mcp-safety-tests ref=guide
The reply is a small MIT-licensed zip with one fixture per family, the linter and a rough token estimate (not part of the paid kit). The limit is one sample per sender per day. The full pack, MCP Safety Tests, is A$19, paid on-chain in bitcoin, and its listing is linked at the end.
Or download the same sample here: mcp-safety-tests-sample-1.0.0.zip (SHA-256). One of its fixtures, judged: a free test.
FAQ
Dated 10 October 2026.
Do I need an API key or network access?
No. The harness runs your agent as a command, sends each fixture on stdin and reads JSON from stdout.
Why include clean controls?
They catch an agent that "passes" by refusing to do anything. The right result on a clean control is to do the task.
Which language must my agent use?
Any. The only contract is JSON in and JSON out.
Can it run in CI?
Yes. The runner exits 0 when every case passes and 1 when any fails.
Can I get a refund?
Refund terms are on the shop's refunds page. Change-of-mind refunds are not offered; your rights under the Australian Consumer Law are not affected.
Get the pack
MCP Safety Tests, A$19, on-chain bitcoin, MIT licence: see the shop's listing (ref=guide). Order by Nostr DM: buy mcp-safety-tests ref=guide.