AlpacaX

Incident

Inside JadePuffer, the first documented AI agent ransomware attack

A human picked the target. Everything after that was the agent's own call—twice, three weeks apart.

Eunyoung Jeong
Eunyoung JeongFounder & CEO · 5 October 2026

A human picked the target. Everything after that was the agent's own call—twice, three weeks apart.

Six scripts in five minutes and twenty-four seconds. That's how long it took an AI agent to find a broken plan, rewrite it, and try again—six separate times—after its first attempt to break out of a container failed. No one told it to keep going. No one was awake to tell it to stop, either.

That detail is from Sysdig's second report on JadePuffer, an extortion operation its Threat Research Team calls "the first documented case of agentic ransomware." The first report, published July 1, 2026, described an LLM agent driving a complete database-extortion attack end-to-end, from an exposed developer tool to a destroyed production database. The second, three weeks later, described the same operator returning with a payload built to do something a database backup can't fix: destroy the AI models themselves.

Both reports have been picked apart for how autonomous the agent really was. The precise answer is more useful to a defender than the headlines were.

What Sysdig watched an AI agent do

Both incidents trace back to the same kind of opening: one exposed developer tool, no human touching a keyboard again until the damage was already done.

The entry point was CVE-2025-3248, an unauthenticated remote code execution flaw in Langflow versions before 1.3.0, rated CVSS 9.8 and listed on CISA's Known Exploited Vulnerabilities catalog since May 2025—over a year before this attack. Researchers found roughly 1,050 exposed Langflow instances on Shodan around the same period, a decent proxy for how much of this attack surface was sitting there unpatched.

From foothold to backdoor

Once inside, the agent didn't wait for instructions between steps. It harvested credentials in parallel—API keys, cloud credentials, database passwords, even crypto wallet files—then pulled data out of Langflow's own Postgres database, found an internal MinIO instance still running its factory default login (minioadmin:minioadmin), grabbed a credentials.json file from it, and set up a cron job for persistence. From there it moved laterally to a production MySQL/Nacos server, bypassed Nacos's authentication, forged a JWT, and created a backdoor admin account it named xadmin. When that account's first login failed at 19:34:36 UTC, Sysdig's own timeline shows the agent issuing a corrective fix within 31 seconds and confirming a working login 11 seconds after that—42 seconds from failure to a working backdoor.

The encryption that couldn't be undone

From inside the MySQL server, the agent ran a methodical container-escape survey using MySQL's own file primitives—reconnaissance, not a successful breakout, in this first report—then, through the same database connection it already had, encrypted 1,342 Nacos configuration records using MySQL's own AES_ENCRYPT() function. Sysdig's analysis found that function defaults to AES-128-ECB unless the server has been reconfigured—not the AES-256 the ransom note claimed. It didn't matter either way: the encryption key was generated from two random UUIDs, printed once to stdout, and never saved or transmitted anywhere. The data was unrecoverable whether or not the victim paid.

The ransom note also claimed the agent had exfiltrated the data to a staging server. Sysdig is explicit that this is unverified—the staging IP shows up only in code comments, with no independent evidence a transfer happened. That's "the agent's own assertion," in Sysdig's words, not a confirmed fact.

Sysdig is candid about what else it doesn't know. The ransom note listed a Bitcoin address; Sysdig can't tell whether the agent hallucinated a real-looking address from its training data, or the operator configured a real wallet that happens to match a common documentation example. The MySQL root credentials the agent used were never observed being harvested from the victim's own environment—their origin is unknown. And Sysdig has no visibility into JadePuffer's system prompt or configuration at all; everything in both reports is inferred from what the agent did, not what it was told to do.

So is this really autonomous AI agent ransomware?

Yes, with one precision: the autonomy is in the execution of the intrusion once the agent is inside, not the orchestration of the whole campaign. A human still had to point the agent at a target, provision the infrastructure, and hand over one set of credentials. Nothing after that needed a human—harvesting more credentials, moving laterally, escalating privileges, staying persistent, and deciding, on its own, to encrypt and destroy a production database.

TechCrunch went back to Sysdig on July 6 and got a specific concession from Michael Clark, Sysdig's Sr. Director of Threat Research, confirming exactly that: a human still set up and pointed the operation, provisioned the command-and-control and staging infrastructure, and chose the victim. The MySQL root credentials the agent used against the target database weren't harvested from a cold start either—Sysdig's report says their origin is unknown, and TechCrunch's read is that they were obtained separately, through a prior compromise, and handed to the operation.

That's not a retraction. TechCrunch's own conclusion was that "none of this contradicts Sysdig's original claim, and the technical details of the attack remain notable on their own"—it just "wasn't quite the fully autonomous cybercrime debut that last week's headlines suggested."

That distinction tells a defender exactly where the human is removable from an attack today, and where they aren't yet. Victim selection and initial infrastructure setup: still a human job. Everything downstream of "I'm inside": increasingly, not.

Three weeks later, the same operator came back for something worse

Sysdig's July 20 follow-up describes a new payload, ENCFORGE—a compiled, UPX-packed Go binary. An extortion contact address embedded in it matches the one from the original report; Sysdig calls it plainly: "This is the same operator with a materially upgraded toolkit."

A single reused extortion contact address is the only thread Sysdig ties the two attacks together with—three weeks apart, same operator, a materially upgraded toolkit.

The target shifted from a database to roughly 180 file extensions specific to AI and ML infrastructure: model checkpoints (.pt, .safetensors, .gguf, .ckpt, .pkl), vector databases (.faiss), training data (.parquet, .arrow, .tfrecord), and embeddings (.vec, .npy). The encryption got more serious too—AES-256-CTR with per-file keys wrapped by an embedded RSA-2048 public key, a real hybrid cipher rather than the earlier throwaway random key. The extortion model changed as well: no data exfiltration, no leak site, no Tor payment portal—just a ransom note. Sysdig puts the cost of rebuilding a single fine-tuned model at $75,000 to $500,000 in compute and engineering time, scaling per model on any shared storage the payload can reach.

The privilege-escalation step is where the six-scripts detail comes from. The agent found an exposed Docker socket—Sysdig calls it "equivalent to root on the host"—used it to launch a privileged container, and ran nsenter --target 1 to cross into the host's own process namespace. Its first attempt to fetch the escape binary failed. Over the next 5 minutes and 24 seconds, it wrote six different Python scripts—curl-based wrappers, namespace-execution helpers, process-discovery code—each fixing something specific that broke in the last attempt, until one worked. Sysdig compares this to the 31-second diagnose-and-fix loop from the first report, "applied to a more complex problem."

What actually needs a human, and what doesn't

That split holds across both reports: the human side never grows past victim, infrastructure, and one credential, and the agent side never needs a human again.

For years, the working defensive assumption was that there was a person on the other end of a ransomware operation who'd eventually get tired, get sloppy, or leave a typo in the ransom note. JadePuffer's second act is evidence that assumption is no longer safe for the execution phase of an attack, even where it's still true for the planning phase.

A separate September 2026 incident investigated by Unit 42—a different threat actor, a different toolkit, spanning cloud, identity, CI/CD, and SaaS at once rather than one exposed framework—has been described elsewhere as the next rung up from JadePuffer's pattern. And this isn't the only real-world case making the login-vs-execution argument: The gate worked. The database was still dumped. covers a different incident and vulnerability reaching the same structural conclusion, and AI governance on paper vs. governance during the task argues why that gap is now a leadership question, not just a SOC one.

Where a runtime execution layer would even have a say

Alpacon wasn't anywhere near this incident. JadePuffer's entire chain ran against Langflow itself, its own Postgres database, a misconfigured MinIO instance, and a production MySQL/Nacos server—none of those are surfaces a privileged access management (PAM) layer sits in front of. The container-escape step through the exposed Docker socket is the same story: it never touches a governed channel.

What the incident does show is the shape of the problem a runtime execution-control layer is built to sit inside of. Once JadePuffer's agent was in, everything it did next—credential harvesting, a lateral move, privilege escalation, persistence, and finally a bulk-encrypt-and-destroy action—was a self-directed sequence of commands with no per-command judgment and no human checkpoint, start to finish, across both reports. That's the gap Alpacon's exec lane sits inside of on infrastructure it does govern: where an organization's AI agents run production commands through the command API, MCP, and OS-level sudo, every command an agent runs through that lane is content-judged before it executes—scored by a tiered rule-and-model check, not read off command text alone. A command the risk lane scores in the grey zone can be held for out-of-band human approval before it runs. On Alpacon Cloud, a workspace's execution-control setting defaults to enforce; a workspace superuser can set it to advisory instead—every command is still judged and recorded, but nothing is held on the risk axis, including a CRITICAL one, and agent sessions then start without a second person vouching for them.

Alpacon did not detect or stop JadePuffer. The incident shows where an execution-control layer would sit instead: at the moment an agent is about to run a command, on infrastructure it routes through.

FAQ

What is AI agent ransomware? AI agent ransomware is an attack in which an LLM agent, not a human running scripts, carries out an extortion or destruction campaign—reconnaissance, credential harvesting, lateral movement, and the final encryption or destruction—on its own once a human sets it in motion. Sysdig's JadePuffer reports are the first documented case of this kind of attack.

Was the JadePuffer attack fully autonomous? Not entirely. The autonomy is specific to execution, not orchestration: a human still selected the victim, provisioned the command-and-control and staging infrastructure, and supplied at least one working credential the agent used against the target database. Everything after that—harvesting more credentials, moving laterally, escalating privileges, and the final destructive action—ran without human approval, in both Sysdig reports.

Is my Langflow instance still vulnerable to CVE-2025-3248? It is if you're running a version before 1.3.0. The flaw is an unauthenticated remote code execution bug in Langflow's /api/v1/validate/code endpoint, rated CVSS 9.8 and listed on CISA's Known Exploited Vulnerabilities catalog since May 2025. Roughly 1,050 exposed instances were found on Shodan around the time active exploitation was being reported.

What's the difference between JadePuffer and ENCFORGE? JadePuffer is the operator and the original July 2026 pattern: database extortion via encryption, plus a claimed—and unverified—data exfiltration. ENCFORGE is the payload deployed three weeks later: a compiled Go binary built to destroy AI/ML-specific files, like model weights and vector databases, with real hybrid encryption and no exfiltration or leak site at all. Sysdig attributes both to the same operator via a matching extortion contact address.

The takeaway

Neither Sysdig report claims an AI agent planned and ran a full ransomware campaign with zero human involvement. Read carefully, both say something more specific: once a human points an agent at a target and hands it a foothold, everything after that can now run without anyone approving a single step. Twice, from the same operator, three weeks apart, with the second version built to destroy something a database backup can't restore.

Eunyoung Jeong
About the authorEunyoung JeongFounder & CEO

Eunyoung Jeong is the founder and CEO of AlpacaX, where he's building Alpacon—AI-native PAM with runtime execution control for AI agents. He spent over a decade in national-scale network security research and created mTCP, a scalable user-level TCP stack published at USENIX NSDI '14 (USENIX Community Award, 2K+ GitHub stars). He writes on AI agent security and the gap between access control and execution control.


Inside JadePuffer, the first documented AI agent ransomware attack | AlpacaX