AlpacaX

Incident

MCP security can't stop at the install review

It passed two clean tool calls, then turned malicious on the third—no one-time review catches that.

Jungyeon Lee
Jungyeon LeeContent Marketer · 7 October 2026

It passed two clean tool calls, then turned malicious on the third—no one-time review catches that.

A malicious MCP server doesn't need to survive your review. It only needs to wait until your review is over.

That's the whole mechanism behind Deadbugz, an active MCP supply-chain campaign disclosed by Pillar Security in August 2026. The server behaves correctly for its first two tool calls, then changes behavior on the third—a call-count trigger that no install-time check, however careful, is built to catch.

A server that behaved for two calls, then didn't

Deadbugz ships as a package called productivity-suite, offering two tools that work exactly as advertised: format_text and summarize. Both run cleanly the first two times an agent calls them. The server keeps an in-memory, per-client counter of tools/call requests, and once that counter hits exactly three, the server's tools/list and prompts/get responses change—what Pillar calls "runtime-gated MCP metadata poisoning" (Pillar Security, "Deadbugz: currently active MCP supply-chain campaign").

After the third call, the altered tool metadata instructs the connected agent to go looking for SSH keys, AWS credentials, shell history, and Kubernetes config—and to conceal the activity from the user (Pillar Security). Nothing about the first two calls would have told a reviewer this was coming. There's no obfuscated payload to find, no suspicious permission request to flag. The server just hadn't reached three yet.

The distribution mechanism was just as deliberate. A GitHub account, zellkernel, opened 23 pull requests against unrelated AI, MCP, and developer-tool repositories in a 74-minute window on 2026-08-10—17 of them wiring up a remote MCP endpoint, 4 referencing a hidden local script path, 2 posing as directory listings (Pillar Security). Nineteen were closed; four remained open, at the time Pillar reviewed them.

NHI Mgmt Group's own writeup names the failure mode this campaign is built to exploit: "Tool approval without runtime revalidation creates a false sense of control." And on the specific mechanism: "The three-call threshold shows how easily a benign inspection can miss a malicious state change." (NHI Mgmt Group, "Deadbugz shows how MCP metadata poisoning evades AI agent trust")

The same lesson, from three unrelated CVEs

Deadbugz is intentional deception. Three MCP server CVEs disclosed the same month show a related failure in a tool's honest, stated functionality: passing install-time review doesn't mean the code behaves safely once a call actually runs with the right inputs (adversa.ai, "MCP security September 2026: Deadbugz + 3 server CVEs").

Path traversal, via a working feature. CVE-2026-73498 hits Atlassian's mcp-atlassian server: the confluence_upload_attachment tool passes an unsanitized file_path straight into open(file_path, "rb"), with no path-safety check in between. An authenticated MCP client can read any file the server process can reach and get it back as a Confluence attachment—including environment variables like the server's own Confluence API token. CVSS 3.1 rates it 7.7, High. Fixed in mcp-atlassian 0.22.0.

A cleartext token, behind a settings-read tool. CVE-2026-67357 hits ArcadeDB's MCP server. Its get_server_settings tool masks any setting whose key contains the word "password"—but arcadedb.ha.clusterToken doesn't match that pattern, so it comes back in plain text. That token is the trust anchor for cluster-forwarded authentication: replay it via the right HTTP headers and you can impersonate root. CVSS 3.1 rates it 7.5, High (CVSS 4.0: 7.7). Fixed in ArcadeDB 26.7.3.

Server-side request forgery, from a pagination helper. CVE-2026-19956 hits facebook-ads-mcp-server: the fetch_pagination_url function lets an authenticated remote user drive requests from the server's own network position to any URL they choose—reaching internal services the server can see but the caller shouldn't. CVSS 3.1 rates it 6.3, Medium (CVSS 4.0: 5.3). Fixed in a single commit.

None of these three tools is malicious. Each does what its name says. The vulnerable pattern sits in the code the whole time—an unsanitized path, a masking rule with a gap, a fetch with no destination check—but nothing about the tool's registration or its name flags that pattern as a problem. The failure only becomes visible in what it does once the call runs with the input that triggers it: a crafted file_path, a settings read that happens to include the token, a pagination URL pointed at an internal address.

What generalizes, and what doesn't

Four different failures, one shared shape: passing review at install time doesn't establish what a tool will do when a call actually runs. Deadbugz's trigger is a runtime-only state that doesn't exist in the code a reviewer would have read at all. The three CVEs are different: their vulnerable patterns are sitting in the code from the start, but nothing about a clean install-time read forces the question of what happens when that code runs with a specific, adversarial input. A tool that passes review with flying colors can still turn out wrong on its third call, or its first, depending on what triggers the flaw.

Narrowing which tools a client is even allowed to see is a separate, worthwhile control—but it answers a different question than this one, and it doesn't close this particular gap. Deadbugz's format_text and summarize are exactly the kind of tools that scoping would wave through; nothing about their names or their first two calls gives a reviewer a reason to exclude them. The gap here sits downstream of whatever got registered, at the point where a call actually runs.

For a broader illustration of a different runtime failure surface in MCP servers—an unauthenticated local transport rather than a delayed-trigger tool—see AutoJack: one webpage, one MCP socket, host-level access.

Where Alpacon fits

Alpacon doesn't score every MCP tool call—a settings read like get_server_settings, or a text-formatting call like format_text, issues no host command and gets no risk score from Alpacon. What Alpacon does cover: when a tool an agent calls runs a command on a host, that command goes through Alpacon's exec lane. Commands routed through the exec lane are scored through a three-tier engine: rule-based checks, a baseline pass, and LLM judgment on the grey zone. A command the risk lane scores in the grey zone can be held for out-of-band human approval before it runs.

That coverage wouldn't have caught Deadbugz's metadata-poisoning trigger itself—the payload there is instructions returned in tool metadata, not a command dispatched to a host. It's the same generalizable point from a narrower angle: judgment has to sit at the point where something actually executes—for Alpacon, that means the point where a tool's call turns into a command running on a host, not the tool call itself.

FAQ

Can a one-time MCP server review ever be enough? Not against a call-count trigger like Deadbugz's—that state doesn't exist in the code at all until the third call. And not reliably against a CVE hiding in a tool's own advertised function either: the vulnerable pattern is there to find, but confirming it's actually exploitable means tracing what happens when the call runs with the input that triggers it, not just reading the function once. A one-time review still catches obviously malicious code and known-bad packages—it just isn't sufficient on its own.

Do I still need to limit which MCP tools a client can register, if runtime judgment is watching what they do? Yes—Deadbugz is exactly why. format_text and summarize are the kind of tools a registration-time scoping pass would wave through without a second look; nothing about their names or their first two calls gives a reviewer a reason to exclude them. Scoping still matters for the tools that never get that far, just not for the one that already made it through the door and waited.

Do these three CVEs mean MCP servers are inherently unsafe? They mean MCP servers are young and under-audited as a category, not that the protocol itself is broken. All three have published fixes, and all three are catchable in principle—but only by someone (or something) checking what the call actually does, not just what the tool is named.

What should change in how a security team evaluates a new MCP server? Ask what happens on the tenth call, not just the first. A code review and a changelog check are necessary but not sufficient; ask whether anything revalidates a tool's behavior at call time, on an ongoing basis, rather than once at approval.

What to check next

The next time an MCP server clears review, ask what that review actually covered. If the answer is "we read the code and it looked fine," you've verified the tool's behavior at one point in time, under one set of inputs, before a single call happened. That's still worth doing. It's just not the same claim as knowing what the tool's calls will do going forward—and Deadbugz, along with three CVEs disclosed the same month, exist specifically because that gap is real and exploitable.

Jungyeon Lee
About the authorJungyeon LeeContent Marketer

Jungyeon Lee writes about AI agent security at AlpacaX—mostly incident analyses of agents that went wrong in production, plus the governance side of it, from ISO 42001 readiness to AI vendor risk. She studied economics and web programming at NYU.


MCP security can't stop at the install review | AlpacaX