AlpacaX

Insights

Why AI governance pilots don't survive scale

The policy wasn't wrong. It just wasn't sized for agents your other teams stood up without telling you.

Eunyoung Jeong
Eunyoung JeongFounder & CEO · September 7, 2026

The policy wasn't wrong. It just wasn't sized for agents your other teams stood up without telling you.

The pilot worked. Three agents, one team, and someone who could point to each one and say what it was allowed to touch. Six months later, the same organization has agents wired into a dozen MCP servers nobody filed a request for, running on tools no one on the original team has heard of, and the governance program that passed review in March can no longer say who's running what. Nothing about the policy changed. What changed is the assumption sitting underneath it: that someone could track every agent. That assumption held at pilot size. It doesn't hold at the size the org grew into.

Gartner predicted in June 2025 that more than 40% of agentic AI projects would be canceled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls. Governance built for a pilot is one of the things dying with the pilots.

Quick answer: Most AI governance tools are sized for the pilot's inventory—a handful of agents one team can name—not the inventory scale produces. A CSA survey commissioned by Token Security found 68% of organizations believe their AI-agent visibility is strong while 82% have discovered previously unknown AI agents in the past year, and Kiteworks' 2026 forecast survey found 60% can't terminate a misbehaving agent. The fix isn't a better inventory—it's a control that doesn't need a current one.

SurveyWhat it measuredFinding
CSA, commissioned by Token SecurityBelief vs. actual AI-agent visibility (418 respondents)68% believe visibility is strong; 82% discovered previously unknown AI agents in the past year
Kiteworks 2026 forecast (225 security/IT/risk leaders)Ability to stop or restrict a running agent60% can't terminate a misbehaving agent; 63% can't enforce purpose limitations
Island MCP-server scan (33,563 builds, 475,865 tools)Risk in the third-party MCP package supply chain40.6% of scanned builds had a tool capable of file/credential access, code execution, or destructive actions
Gartner, June 2025Reasons agentic AI projects get canceledPredicts over 40% of agentic AI projects canceled by end of 2027, citing cost, unclear value, or inadequate risk controls

Why does a pilot's governance look complete when it isn't?

TL;DR: Confidence in AI-agent visibility and actual visibility are not the same measurement, and the gap between them is the whole story. A CSA survey commissioned by Token Security found 68% of organizations believe they have strong visibility into their AI agents. In the same survey, 82% discovered previously unknown AI agents in the past year.

At pilot size, "we have visibility" is true because visibility barely takes any work. Three agents, one Slack channel, one person who wrote the request tickets. Ask that person what's running and they can tell you, correctly, in about ten seconds.

That's the exact model that stops working once a second team spins up its own agent, a third team wires one into an MCP server the first two never heard of, and a fourth starts an evaluation nobody's tracked in the governance doc. The org's confidence in its own visibility doesn't drop when this happens—it usually stays high, because nobody re-measures it. The CSA survey, commissioned by Token Security, found both numbers on the same 418 respondents: 68% believe their AI-agent visibility is strong, and 82% discovered previously unknown AI agents in the past year, and 41% of all respondents have found one more than once. Add one more data point from the same report: only 21% have a formal decommissioning process—which suggests that when a pilot ends, whatever access it had doesn't reliably get revoked either. The report's own framing is close to a thesis statement for what comes next: "as agents gain greater autonomy, governance must evolve into a more unified, operational model that can sustain control at scale."

What does the breakdown actually look like once it happens?

TL;DR: Kiteworks' 2026 forecast numbers show what a governance program that outgrew its own coverage looks like in practice—60% can't terminate a misbehaving agent, 63% can't enforce purpose limitations. A related pattern shows up in a 2026 MCP-server scan: in 40.6% of scanned builds, at least one tool appeared capable of accessing files or credentials, executing code, or taking destructive actions.

Three months into a pilot, the answer to "can you stop this agent if it does something wrong" is yes, because someone built the kill switch for exactly those three agents. A year later, in Kiteworks' 2026 forecast survey of 225 security, IT, and risk leaders, 60% cannot terminate a misbehaving agent, and 63% cannot enforce purpose limitations on what those agents are authorized to do.

A related pattern shows up in a 2026 MCP-server security scan by Island, a browser-security vendor, that looked at the supply side of the same problem: 33,563 published MCP server builds containing 475,865 tools. 49% produced at least one non-informational finding once benign inventory markers were excluded, and in 40.6% of builds, at least one tool appeared capable of accessing files or credentials, executing code, or taking destructive actions. Among packages with a known maintainer count, 84% listed a single maintainer or publisher identity—which Island notes may be a person, a company, or an automated publishing account—and among those with an identifiable owner, 92% had no match to a GitHub-verified organization: an attribution gap in the public package supply chain that most third-party MCP servers are pulled from.

Why doesn't a correct policy survive the move from pilot to scale?

TL;DR: A governance program scoped and tested against a small, known inventory has no built-in way to extend to the next team's MCP server or the tool nobody filed a request for—coverage that was complete on day one silently stops being complete as the inventory grows, independent of whether the policy itself was ever wrong.

Usually, nobody wrote a bad policy. The policy that governed three agents on one team was probably fine for three agents on one team. What it didn't have—because nothing forced anyone to build it—was a mechanism that automatically extended to the fourth team, the new MCP server, the tool evaluation that started without a ticket. Coverage that was complete when the inventory was small and known has no property that keeps it complete once the inventory grows and stops being known. That's a structural gap, not a policy error, and it explains why "the policy was fine at review" and "the org has no idea what's running six months later" can both be true at once.

What's actually at stake if this doesn't get fixed?

TL;DR: Gartner's 40%-cancellation prediction and Kiteworks' 2026 forecast numbers are two independent surveys pointing at the same failure—one names the technical gap, the other names it a reason projects get killed.

Put the two data sets side by side and they point at the same failure without needing to share a sample or a population. Kiteworks' numbers are the technical reality: 60% can't terminate a misbehaving agent, 63% can't enforce purpose limits. Gartner names inadequate risk controls as one of three reasons agentic AI projects get canceled, alongside escalating costs and unclear business value. A program that can't say what's running, or stop it if it misbehaves, fits that description—whether or not anyone on the team would use that phrase for their own work.

What does governance that actually survives scale need?

TL;DR: A control that depends on someone maintaining a complete, current map of every agent, tool, and team breaks the moment that map goes stale. The more durable answer doesn't need a current inventory to judge what reaches it—it needs the traffic routed through it in the first place.

The usual answer is a better map: a more complete inventory, a stricter onboarding process for new agents, a mandatory registration step before any team can stand up an MCP server. Some of that is genuinely useful. None of it solves the structural problem, because the map is only ever as current as the last time someone updated it, and the whole reason pilots stop being pilots is that teams move faster than the update cycle.

The alternative is a control that doesn't need a current per-agent inventory to judge what reaches it—it needs the traffic routed through it. That's the argument for putting governance at the execution layer—the layer Alpacon, an AI-native PAM, is built on: identity decides who connects, policy decides what's allowed on paper, and execution control is where each command is judged before it runs. A Work Session is opened with an identity, a declared intent, and a scope ceiling, and commands in the exec lane are judged against that declared intent, not just against command text—Alpacon can hard-deny a command or route it to a human before it executes—enforcement is the default, with a monitor-and-record mode for teams that want to watch the judgments first. The judgment is per-command—it doesn't depend on anyone maintaining a current list of agents. Getting a new team's agents onto that path is its own onboarding step; finding the MCP server nobody registered in the first place is still a separate job.

The takeaway

Your governance program didn't fail because someone wrote the wrong policy. It failed because it was built for a headcount and an inventory that no longer exist, and nothing in it noticed when that stopped being true. The fix isn't a better version of the same map—it's making the path itself the requirement, so a command on that path is judged on its own terms.

If you're scaling AI agents past the size where someone can track them by name, that's the layer worth building before the pilot outgrows its own governance, not after.

Eunyoung Jeong
About the authorEunyoung JeongFounder & CEO

Eunyoung Jeong is the founder and CEO of AlpacaX, where he's building Alpacon—AI-native PAM with runtime execution control for AI agents. He spent over a decade in national-scale network security research and created mTCP, a scalable user-level TCP stack published at USENIX NSDI '14 (USENIX Community Award, 2K+ GitHub stars). He writes on AI agent security and the gap between access control and execution control.


Why AI governance pilots don't survive scale | AlpacaX