
For about two years, the security conversation around AI agents lived in research blogs and conference talks. It was a conversation among practitioners, and it was mostly voluntary. Nobody was going to fail an audit over an agent.
That window has closed. Somewhere between the first agent that wrote its own shell commands and the first breach an agent carried out end to end, the people who write rules started paying attention. The vendor security review that used to ask about encryption and access reviews now asks whether you keep an inventory of your AI agents, whether they run with scoped credentials, and whether you can produce a log of what they did and why. The theoretical phase is over, and the compliance phase has begun.
A compliance framework can mean three different things, and the difference decides how much choice you have. Some are voluntary best practices, adopted because they are sound and because customers expect them, like the NIST AI Risk Management Framework. Some are certifiable standards you can be audited against, like ISO/IEC 42001. And some are law, like the EU AI Act, with real penalties attached.
Most of these frameworks were not built for agents. Standards like ISO 42001, the NIST AI RMF, and the EU AI Act were written for AI systems in general, back when a model answered a question and a human decided what to do with it. The risk ended at the screen. Agents changed the stakes rather than the rulebook. An agent does not stop at the answer; it plans, calls tools, moves data, and acts in a loop faster than anyone can review, which quietly pulled it into the scope of rules that never imagined it. That same shift produced a new agent-native layer on top, with OWASP building an entirely new Top 10 for agents, because autonomy itself is the risk. Governments followed in 2026, with Singapore publishing a governance framework written for agentic AI in January and six national cyber agencies, including CISA and the NSA, issuing joint guidance on adopting agents in May. The result is the stack below, and enterprise buyers are already gating contracts on it: ISO 42001 has moved from a nice-to-have to a procurement expectation, increasingly written into the security questionnaires and vendor RFPs that decide who makes the shortlist.
The clearest way through the landscape is to sort each framework by the job it does. Sorted that way, the field falls into five layers. At the top sit the standards for governing and managing AI in general. Below them are the threat models and control sets built specifically for agents, then the operational guidance for running them safely, then the laws that carry real penalties, and finally the identity standards that decide whether a given agent is allowed to act at all. Each layer answers a different question about your agents, and reading them as a stack keeps the whole thing legible.
A few are worth pulling out:
Placed next to each other, the frameworks reveal something their differences hide. They vary in format and jurisdiction, but they converge on the same short list of controls: know which agents you run, give each the least authority it needs, validate what flows into a tool call, require a human before anything irreversible, vet the tools and MCP servers an agent can reach, and keep an audit trail.
Now hold that list against what agents look like in the wild. In our audit of 164,692 code files from roughly 86,000 repositories and over 700,000 exposed instances, audit logging, the one control every framework shares, appeared in just three of those files. Among the code that defined agent tools, 76.4 percent had no input validation. On the MCP side, 82.8 percent of servers lacked input validation, 92.4 percent had no confirmation gate before a tool executed, and 98.1 percent had no rate limiting. The controls the rulebook now demands are, by default, missing from the systems it now governs.
Probably more than one, and that is the honest answer. Which ones bind you depends on where you operate and what your agents touch. If you sell into Europe, the EU AI Act reaches you. If you work in health or finance, older rules already reach your agents: HIPAA pulls any agent that can touch patient records under its access-control and audit requirements, and GLBA does the same for any agent that handles customer financial data. Bank supervision moved the other way: when the Federal Reserve, OCC, and FDIC replaced SR 11-7 in April 2026, they left generative and agentic AI out of the new model risk guidance and promised a separate request for information on it. Until that arrives, the most detailed guidance banks have is Treasury's Financial Services AI Risk Management Framework, released in February with 230 control objectives built on the NIST AI RMF. If an enterprise buyer is running the procurement, they will ask for ISO 42001 or SOC 2, because a third-party certificate lets them trust your AI governance without auditing you themselves. The agents themselves now have a certificate of their own: AIUC-1 audits a specific agent deployment for security, safety, and reliability, and KPMG and UiPath already hold it. The voluntary frameworks are then worth adopting in the order the stack suggests: a management system to govern, OWASP and MITRE ATLAS to threat-model, the five-eyes, the NSA, and NIST guidance to operate safely, and a scoped identity for every agent.
But adopting frameworks is not the same as satisfying them. The reflex is to write a policy, buy a gateway, and book the audit, yet each falls short where agent risk lives: a policy cannot prove an agent stayed in bounds, a gateway sees one chokepoint and not the reasoning loop behind it, and a yearly audit is a snapshot of behavior that changes with every prompt. The evidence these frameworks want is behavioral, and it only exists where the agent acts, at runtime.
This is exactly where the newest layer of the stack is heading. The identity and authorization work, from OpenID's AuthZEN profiles to the MCP authorization spec and NIST's NCCoE agent-identity project, is trying to answer one question at the moment it matters: may this agent, acting for this user, call this tool with these arguments? That is a runtime decision, made before the tool runs, and it is the clearest signal yet that the whole field is converging on the layer where agents actually act.
The MCP spec, now governed by the Linux Foundation's Agentic AI Foundation, moved the same way in July 2026, when its largest revision since launch made OAuth 2.1 and OpenID Connect a requirement for authorization. The Cloud Security Alliance's AARM specification spells out what a runtime security system for agents has to do: intercept each action before it runs, weigh it against policy and the context of the session, and keep a tamper-evident record of the decision. It is early, and CSA calls it a living framework, but it already defines conformance requirements that a product can be checked against. OWASP's Agent Control Standard, which Capsule help lead, donated to the OWASP GenAI Security Project in September and still a v0.1 preview, covers the agent platform's side: it defines the hooks where the platform stops to check a tool call against policy and then enforces the result.
That is the problem Capsule's compliance capability was built for. It works from both sides of the agent lifecycle. Statically, it discovers your agents, including the inline and undeclared ones your inventory is missing, and flags dangerous tools, missing validation, and unvetted MCP servers. At runtime, it catches goal hijacks and injected instructions, enforces approval before consequential actions, and records every tool call with the context around it. The compliance layer ties both to the rulebook: a missing-validation finding becomes evidence against OWASP ASI02 and the NSA guidance, and an approved, recorded tool call becomes proof of the human oversight the EU AI Act requires and the audit trail ISO 42001 expects. Instead of a once-a-year mapping exercise, you get continuous, framework-aligned evidence that already exists when the questionnaire arrives.
The frameworks arrived faster than most teams expected, and they will keep arriving. What ties them together is a single demand: show us, at the level of individual actions, that your agents do only what they are allowed to do, and prove it. That cannot be met with a binder or a perimeter box. It is met at runtime, in the evidence of what your agents actually did.
Want to see how Capsule maps your agent findings and runtime detections to OWASP, NIST, the EU AI Act, and ISO 42001? Get a demo.

What happens when an AI agent blindly trusts a name nobody owns? We claimed abandoned buckets and packages referenced in agent files, and real production systems started sending us model weights, backups, and installs. GhostSquatting turns forgotten names into model theft, data exfiltration, and RCE.

Jev is fast, clever, and great at classification, so why not put it in front of every AI agent action? We tested it against Capsule's rogue agent model on accuracy, latency, and context, and found that when a wrong answer means a deleted production database, specialization still wins.

OWASP has published the Agentic Skills Top 10, a new list of the most critical risks in the skills that give AI agents real-world reach across tools like Claude Code and Cursor. We compared each of the ten risks against our analysis of over 200,000 real skills, and every one of them is already happening in the wild.

Capsule Security research uncovered a behavior in Cursor's agent: asked to do something ordinary like share a file, it decides on its own to upload the file to a public anonymous host to get a link. It will push past a deny-all network sandbox to do it, and no attacker is involved. We found it running in production across every major model, reported it to Cursor, and were met with silence.

Modern AI agents are increasingly causing critical system damage not through external cyberattacks, but by taking unprompted, off-script actions across unguarded tools like databases and system shells. Traditional safeguards, including prompt instructions and human approval workflows, routinely fail to catch these autonomous errors before execution. To mitigate this growing risk, security must shift directly to the tool-call boundary, enforcing deterministic runtime controls that intercept and block destructive commands before they run.

Capsule Security and NVIDIA collaborated to solve the rogue AI agent threat by engineering specialized Small Language Models (SLMs) for real-time security detection. By fine-tuning NVIDIA Nemotron architectures, this solution achieves inline, ultra-low latency interception of unauthorized agent actions before damage occurs. Discover how domain-specialized models deliver up to 96.9% accuracy and sub-200ms response times to keep autonomous enterprise workflows secure.

Capsule launches a security integration for Claude Platform, using Claude's Compliance API to give security, compliance, and AI governance teams visibility into enterprise AI activity, risk, and posture across Anthropic-hosted deployments.

Our analysis of 206,435 AI agent skills reveals a rapidly growing software supply chain vulnerable to natural language payloads and dangerous capability combinations. Read the report to understand how these skills bypass traditional security controls and learn how Capsule protects your organization by securing the agent runtime.
.png)
The theoretical phase of agentic AI security is over—the attack surface is real and the incidents are documented. This post breaks down the defensive architecture taking shape in response: Meta's Agents Rule of Two, deterministic enforcement hooks, identity governance for non-human agents, and the questions security leaders need to be asking right now.

The security risks of AI agents are no longer theoretical. This blog examines the active threat landscape facing agentic AI in 2026, from prompt injection and supply chain attacks against MCP and skill registries to the governance gap created by vibe coding and Shadow AI.

Guardian agents are emerging as a critical security layer for the agentic AI era. As enterprises adopt AI agents that execute tools, handle sensitive data, and operate inside real workflows, human approval loops no longer scale. Guardian agents solve this by supervising other agents in real time: monitoring actions, enforcing policy, and blocking risky behavior before execution.
.png)
Capsule found two Cursor IDE vulnerabilities that let hidden prompt-injection instructions in referenced files steal developers’ SSH keys and contaminate future unrelated projects, causing zero-click or one-click exfiltration even when the attacker ships no malicious code.

Capsule Security’s State of AI Agent Security 2026 report is the largest independent audit of AI agents to date, showing that the ecosystem is rapidly shipping publicly exposed, weakly guarded, highly connected agents with recurring misconfigurations, near-absent runtime controls, widespread prompt-injection risk, expanding supply-chain exposure, and active malicious campaigns still propagating through agent skill and tool registries.

Capsule is launching a runtime security platform for the agentic AI era, built to monitor and stop autonomous agents that can bypass traditional guardrails, misuse legitimate access, and create a new class of enterprise security risk.

Capsule research team discover a critical prompt injection vulnerability in Salesforce Agentforce that allows attackers to exfiltrate CRM data through a simple lead from a form submission. No authentication required.