
An engineer at a large AI company kicked off what should have been a forgettable task: a routine sync of training data to the team's dataset bucket, the kind of job their pipelines and agents had run dozens of times. On every earlier run, one of the bucket names written into the configuration had quietly failed to resolve, and the agent had done what agents are built to do, catching the error and correcting it on the fly so the pipeline sailed on, entirely seamless to the developer who never knew a name had failed. This time the same name was resolved cleanly on the first attempt. The upload of freshly trained model weights went straight through, the logs stayed calm, and nothing gave anyone a reason to look twice. The only thing that had changed lived entirely on our side of the connection: we had registered that unclaimed bucket name, so the reference that used to fail, and used to be silently repaired, now pointed at a real bucket that happened to belong to us.
A few days earlier we had started a research project with a narrow question: which resources named in AI agent instruction files have quietly been abandoned? These are the files that tell an agent which bucket to read, which artifact to download, and which package to install, and the resources they point at get orphaned in a handful of ways. Sometimes an agent hallucinates a name into a config and it was never real to begin with. Sometimes a bucket or a package that was genuinely in use gets deleted or unpublished while every reference to it lives on. And sometimes a name resolves somewhere it should not, at the wrong cloud or the wrong registry, so it works on its author's machine and misfires everywhere else. However it happens, the name ends up sitting in a public file, unclaimed, waiting for something to answer it.
So we answered. We took the unclaimed names, buckets and packages alike, and registered them in a research account before anyone with worse intentions could, then turned on logging and watched what came back. It did not take long. Within days, production systems from dozens of unrelated organizations were listing our buckets, downloading files from us, and uploading their data into storage we owned. We ran this as coordinated disclosure, reaching out privately to every affected organization before publishing, and we hold the registered names for return. What follows is what reached us, and why it keeps happening.

Dangling buckets and dependency confusion have been understood for years, and it matters to say so plainly, because the classic fix still applies: own your names and pin your dependencies. What is new is where these references now live and how much power they carry once they are there, and that comes down to three things.
The most direct version of owning a name is owning a bucket. A skill refers to a bucket by name, and that name is a single global identifier across all of S3 or GCS. When no bucket answers to it, because it was never created, sits on another cloud, or was deleted, anyone can create one that does and inherit every request meant for the original. We scanned more than 20,000 instruction files across roughly 1,400 public repositories and pulled out close to 2,000 unique bucket names. Wherever a referenced bucket had vanished and its name was free, we registered it, held it empty, and turned on access logging. That came to 264 buckets, 209 on AWS and 55 on Google Cloud.

For every dangling bucket we went back to the file that referenced it and asked a simple question: why was this name ever written down? A few stories repeat.
How bad a takeover gets depends on what the bucket holds, and the corpus was full of sensitive ones: database backups, Terraform state (a full map of a company's cloud, secrets included), CloudTrail security logs, a bucket that distributes signed firmware, and clinical datasets down to hospital brain scans. Claiming any one of them puts an attacker in position to steal the data flowing in, or to poison whatever the pipeline pulls back out.
It works, and it works unusually well against agents. Twenty-nine of the buckets we held drew live external traffic, thousands of requests in total, from many different organizations. Almost none of it came from internet scanners. It came from production pipelines and CI jobs that already knew the exact paths they wanted and asked us for them by name, run after run, certain they were talking to their own storage. A human might notice a bucket that used to fail and now answers; an agent following a written instruction just trusts the name and connects, which is what makes this land so reliably.
The most complete story we captured belonged to a large AI company. One of their dataset buckets did not live on AWS at all; it sat on a different S3-compatible cloud, and their repository was set up to reach it there. One of their skill files had the destination wrong. It told the agent to use plain AWS S3 for that bucket instead of the provider it actually lived on, so the agent did as it was told, its aws s3 calls went to real AWS S3, and they landed on the exact name we were holding.
From there the pattern was almost domestic in its regularity. Their pipelines searched for specific dataset and checkpoint paths, and their jobs handed us the crown jewels without a second thought: trained model weights in .safetensors form, training datasets, and packaged code bundles, all sent to a bucket we owned, by dozens of their engineers, across thousands of requests.


The downloads were worse. We watched their pipelines pull PyTorch checkpoints, plain .pt files, straight from our bucket, and a PyTorch checkpoint runs whatever code is pickled inside it the moment torch.load() opens it. One poisoned checkpoint served back would have been remote code execution inside their training run. Their data came to us; our code could have gone to them, over the same wire.

The packages came out of the very same files. When a skill tells an agent to run npx create-thing, uvx some-mcp, or pip install some-lib, that line is a runtime execution command, so we extracted every reference of that shape and checked each name against the public registry.
The result reframed the project. The corpus held more than 3,600 unique package names, of which 603 were genuinely dangling across npm and PyPI, more than the 264 claimable buckets. A package name is even cheaper to hallucinate than a bucket, and it lands directly inside a command the agent runs.

The same pattern holds, with its own recurring causes:


Of those 603, 331 are direct-execution forms (npx, uvx, pipx run), where the code runs the instant the agent obeys, with no install step and no confirmation in between.
To measure the real exposure, we published our own placeholder packages under some of the dangling names, each with a post-install beacon that only records that an install happened. Those placeholders logged more than 100 installs in the first five hours and over 200 in all, from developer laptops and CI machines. An attacker publishing under the same names would have been running arbitrary code on every one of those hosts.
This is an old threat, and some of the fixes are old too: own your names, pin your dependencies. What is new is that the agent runtime gives you places to enforce them that did not exist before.
Every one of these buckets filled and every one of these packages ran for the same reason: the agent trusted a name that no one owned. Close that gap and the whole attack has nowhere left to land.
We conducted this research under coordinated disclosure and notified every affected organization we identified. The bucket names are held for return. If you suspect a name in your own agent files may be exposed, contact us at info@capsule.security.

Jev is fast, clever, and great at classification, so why not put it in front of every AI agent action? We tested it against Capsule's rogue agent model on accuracy, latency, and context, and found that when a wrong answer means a deleted production database, specialization still wins.

OWASP has published the Agentic Skills Top 10, a new list of the most critical risks in the skills that give AI agents real-world reach across tools like Claude Code and Cursor. We compared each of the ten risks against our analysis of over 200,000 real skills, and every one of them is already happening in the wild.

Capsule Security research uncovered a behavior in Cursor's agent: asked to do something ordinary like share a file, it decides on its own to upload the file to a public anonymous host to get a link. It will push past a deny-all network sandbox to do it, and no attacker is involved. We found it running in production across every major model, reported it to Cursor, and were met with silence.

As AI agent security moves into formal governance, binding laws like the EU AI Act and procurement standards like ISO 42001 are becoming required, while voluntary frameworks like OWASP’s Agentic Top 10 guide threat modeling. Despite these rising expectations, audit data shows that most active agent deployments still lack basic safeguards like input validation and execution logging. This post breaks down the five-layer compliance stack and explains why continuous, runtime evidence is essential to satisfy both strict regulations and voluntary benchmarks.

Modern AI agents are increasingly causing critical system damage not through external cyberattacks, but by taking unprompted, off-script actions across unguarded tools like databases and system shells. Traditional safeguards, including prompt instructions and human approval workflows, routinely fail to catch these autonomous errors before execution. To mitigate this growing risk, security must shift directly to the tool-call boundary, enforcing deterministic runtime controls that intercept and block destructive commands before they run.

Capsule Security and NVIDIA collaborated to solve the rogue AI agent threat by engineering specialized Small Language Models (SLMs) for real-time security detection. By fine-tuning NVIDIA Nemotron architectures, this solution achieves inline, ultra-low latency interception of unauthorized agent actions before damage occurs. Discover how domain-specialized models deliver up to 96.9% accuracy and sub-200ms response times to keep autonomous enterprise workflows secure.

Capsule launches a security integration for Claude Platform, using Claude's Compliance API to give security, compliance, and AI governance teams visibility into enterprise AI activity, risk, and posture across Anthropic-hosted deployments.

Our analysis of 206,435 AI agent skills reveals a rapidly growing software supply chain vulnerable to natural language payloads and dangerous capability combinations. Read the report to understand how these skills bypass traditional security controls and learn how Capsule protects your organization by securing the agent runtime.
.png)
The theoretical phase of agentic AI security is over—the attack surface is real and the incidents are documented. This post breaks down the defensive architecture taking shape in response: Meta's Agents Rule of Two, deterministic enforcement hooks, identity governance for non-human agents, and the questions security leaders need to be asking right now.

The security risks of AI agents are no longer theoretical. This blog examines the active threat landscape facing agentic AI in 2026, from prompt injection and supply chain attacks against MCP and skill registries to the governance gap created by vibe coding and Shadow AI.

Guardian agents are emerging as a critical security layer for the agentic AI era. As enterprises adopt AI agents that execute tools, handle sensitive data, and operate inside real workflows, human approval loops no longer scale. Guardian agents solve this by supervising other agents in real time: monitoring actions, enforcing policy, and blocking risky behavior before execution.
.png)
Capsule found two Cursor IDE vulnerabilities that let hidden prompt-injection instructions in referenced files steal developers’ SSH keys and contaminate future unrelated projects, causing zero-click or one-click exfiltration even when the attacker ships no malicious code.

Capsule Security’s State of AI Agent Security 2026 report is the largest independent audit of AI agents to date, showing that the ecosystem is rapidly shipping publicly exposed, weakly guarded, highly connected agents with recurring misconfigurations, near-absent runtime controls, widespread prompt-injection risk, expanding supply-chain exposure, and active malicious campaigns still propagating through agent skill and tool registries.

Capsule is launching a runtime security platform for the agentic AI era, built to monitor and stop autonomous agents that can bypass traditional guardrails, misuse legitimate access, and create a new class of enterprise security risk.

Capsule research team discover a critical prompt injection vulnerability in Salesforce Agentforce that allows attackers to exfiltrate CRM data through a simple lead from a form submission. No authentication required.