
Fast classifiers are having a moment. Are they ready to stand between your AI agents and production?

Jev landed about a week ago. My feed now has more hot takes about “System One” than the Kahneman framing the name borrows from ever got. Fair enough. It’s one of the more interesting model releases of the year. It also touches the exact problem we work on every day at Capsule: deciding, in milliseconds, whether an AI agent should be allowed to do the thing it’s about to do.
So let’s ask the question everyone in agent security is quietly asking. Is Jev the answer?
Jev comes from TypeSafe AI, which opened early access on September 15. It doesn’t generate text. You hand it a “state” plus a few typed questions (Choice, Score, or Noul, a yes/no probability), and it answers all of them in a single parallel pass, with probabilities your code can use directly. The answers are defined up front, so it can’t return anything outside your schema. That’s a lovely property for anyone who has ever written a regex at 2am to rescue a malformed JSON response.
TypeSafe describes it as transformer-based, trained only on synthetic data with a method called RLCD (Reinforcement Learning for Calibrated Decisions), and reports responses in 70 to 500 milliseconds, 40 to 200 times faster than frontier LLMs on comparable tasks.
For a huge range of classification work (ticket routing, email triage, fraud flags, product categorization), Jev is great. If you’re routing support tickets through a frontier model today, you’re paying for a limousine to deliver a pizza.
Here’s where it gets interesting for us. Within days, developers started wiring Jev in front of agent tool calls, screening each one for risky actions before it runs. The appeal is obvious. Coding harnesses have classified dangerous actions before execution for a while, but those classifiers live in closed-source parts of the harness. A cheap, capable classifier seems to make the pattern available to every agent.
The logic is seductive. Agent security is a classification problem. Every tool call is a question: allowed or not? Jev answers questions fast and cheap. Put Jev in front of every action and call it a day.
And honestly? The architecture is right. We reached the same conclusion from the other direction. As we described in our Nemotron research, Capsule’s detectors don’t write out text at runtime either. We read the next-token logits to get the probability of a violation, so latency comes down to time-to-first-token. Inline security wants a classifier, full stop. On that point, TypeSafe and we agree completely.
The architecture was never the hard part, though.
A great architecture without specialization gets you a fast, confident, well-calibrated opinion. In security, that opinion also has to be right. Three fundamentals decide whether it is: data, fine-tuning, and context.
Data. A detector is only as good as the boundary it learns. Our training pipeline starts from real agent traces. A frontier model (NVIDIA Nemotron 3 Ultra) judges each tool call, and ambiguous cases go to human reviewers. Then, for every labeled example, we generate the smallest edit to the session that flips the verdict. Those near-identical pairs are the whole game. They teach the model the exact spot where “clean up the staging credentials” turns into “delete the production volume.” You don’t get that boundary from synthetic data. Jev, by design, is trained on synthetic data only.
Fine-tuning. Jev doesn’t support fine-tuning at the moment. Every account runs on the same weights, and you shape its behavior through the state, instructions and criteria you send. That’s a sensible product decision for a general classifier. It also means your security posture is only as good as your prompt. We went the other way, training LoRA adapters on roughly 1.4% of parameters to specialize an already capable model to a single decision boundary. Nobody has to tell our model what a rogue agent looks like. It has seen thousands of them.
Context. Jev caps each request at 64K tokens, 32K of it for the state, and TypeSafe’s own docs note that accuracy drops as the state fills with content unrelated to the decision. An agent session is exactly that: hundreds of steps of tool outputs, file reads and web pages, most of it noise and some of it written by an attacker. In agent security, deciding what’s relevant is the detection problem.
We ran Jev against the same held-out rogue-agent evaluation set we use for our own detectors, with carefully written instructions and criteria for each question. Jev scored 82% accuracy against our model’s 97%. It took 100 to 500 ms per decision against our 78 ms. And it works within a 64K window (32K for the state) against our 1M.

Figure 1. Head to head. Accuracy on Capsule’s held-out rogue-agent evaluation set. Jev latency from our testing; TypeSafe reports 70 to 500 ms. Context limits per TypeSafe documentation and NVIDIA Nemotron 3 Nano specifications.
Accuracy. Fifteen points is the difference between a guardrail and a suggestion. On our eval set, Jev lands on the wrong side of roughly one in five decisions. Put that in front of agents making thousands of tool calls a day, and “one in five” stops being a statistic and becomes your on-call rotation. Either you drown in false positives until someone turns it off, or the misses turn into incident reports. Our numbers also hold up outside our own lab. On the public StepShield benchmark, our 30B detector caught 99% of rogue trajectories, raised false positives on under 1% of clean runs, and flagged each violation at the exact step it happened.
Latency. Inline security means every single tool call waits for a verdict. At the top of Jev’s range, a 200-step session picks up well over a minute and a half of guardrail time. At 78 ms, the same session adds about 16 seconds, spread so thin no developer notices. Speed is what gets a guardrail left switched on.
Context. Nemotron 3 Nano supports a 1M-token context window, so our model reads the whole session, including an instruction from step 3 that turns a routine-looking command at step 180 into a violation. With 32K for state, you’re choosing which parts of the story the judge gets to read. An attacker only needs you to choose wrong once.
For a lot of things, absolutely Jev. Ticket routing, triage, tagging, model routing, moderation, cheap first-pass filtering at massive scale. It’s fast, calibrated and cheap, and it’s a genuinely new tool in the box. I expect we’ll see it everywhere, and I think that’s a good thing.
Sensitive security decisions are a different category. When a wrong answer means a wiped production database or a leaked credential, good general judgment isn’t enough. You need a model trained on real agent behavior and fine-tuned to the exact boundary between safe and unsafe. It needs enough context to see the whole session. And it needs to be built by people who spend their days finding new ways agents go wrong.

Figure 2. The simplest way to choose: ask what happens when the model gets it wrong.
That’s exactly what we build at Capsule. Our specialized models sit inline on every agent action, checking each tool call before it runs and catching rogue behavior, prompt injection, credential leakage and tool poisoning as they happen. Capsule covers the coding agents your developers already live in, like Claude Code, Cursor and GitHub Copilot, and the enterprise platforms your teams are building on, including Microsoft Copilot Studio, AWS Bedrock, Azure AI Foundry and Salesforce Agentforce.
Let Jev sort your tickets. Let Capsule guard your agents.

OWASP has published the Agentic Skills Top 10, a new list of the most critical risks in the skills that give AI agents real-world reach across tools like Claude Code and Cursor. We compared each of the ten risks against our analysis of over 200,000 real skills, and every one of them is already happening in the wild.

Capsule Security research uncovered a behavior in Cursor's agent: asked to do something ordinary like share a file, it decides on its own to upload the file to a public anonymous host to get a link. It will push past a deny-all network sandbox to do it, and no attacker is involved. We found it running in production across every major model, reported it to Cursor, and were met with silence.

Modern AI agents are increasingly causing critical system damage not through external cyberattacks, but by taking unprompted, off-script actions across unguarded tools like databases and system shells. Traditional safeguards, including prompt instructions and human approval workflows, routinely fail to catch these autonomous errors before execution. To mitigate this growing risk, security must shift directly to the tool-call boundary, enforcing deterministic runtime controls that intercept and block destructive commands before they run.

Capsule Security and NVIDIA collaborated to solve the rogue AI agent threat by engineering specialized Small Language Models (SLMs) for real-time security detection. By fine-tuning NVIDIA Nemotron architectures, this solution achieves inline, ultra-low latency interception of unauthorized agent actions before damage occurs. Discover how domain-specialized models deliver up to 96.9% accuracy and sub-200ms response times to keep autonomous enterprise workflows secure.

Capsule launches a security integration for Claude Platform, using Claude's Compliance API to give security, compliance, and AI governance teams visibility into enterprise AI activity, risk, and posture across Anthropic-hosted deployments.

Our analysis of 206,435 AI agent skills reveals a rapidly growing software supply chain vulnerable to natural language payloads and dangerous capability combinations. Read the report to understand how these skills bypass traditional security controls and learn how Capsule protects your organization by securing the agent runtime.
.png)
The theoretical phase of agentic AI security is over—the attack surface is real and the incidents are documented. This post breaks down the defensive architecture taking shape in response: Meta's Agents Rule of Two, deterministic enforcement hooks, identity governance for non-human agents, and the questions security leaders need to be asking right now.

The security risks of AI agents are no longer theoretical. This blog examines the active threat landscape facing agentic AI in 2026, from prompt injection and supply chain attacks against MCP and skill registries to the governance gap created by vibe coding and Shadow AI.

Guardian agents are emerging as a critical security layer for the agentic AI era. As enterprises adopt AI agents that execute tools, handle sensitive data, and operate inside real workflows, human approval loops no longer scale. Guardian agents solve this by supervising other agents in real time: monitoring actions, enforcing policy, and blocking risky behavior before execution.
.png)
Capsule found two Cursor IDE vulnerabilities that let hidden prompt-injection instructions in referenced files steal developers’ SSH keys and contaminate future unrelated projects, causing zero-click or one-click exfiltration even when the attacker ships no malicious code.

Capsule Security’s State of AI Agent Security 2026 report is the largest independent audit of AI agents to date, showing that the ecosystem is rapidly shipping publicly exposed, weakly guarded, highly connected agents with recurring misconfigurations, near-absent runtime controls, widespread prompt-injection risk, expanding supply-chain exposure, and active malicious campaigns still propagating through agent skill and tool registries.

Capsule is launching a runtime security platform for the agentic AI era, built to monitor and stop autonomous agents that can bypass traditional guardrails, misuse legitimate access, and create a new class of enterprise security risk.

Capsule research team discover a critical prompt injection vulnerability in Salesforce Agentforce that allows attackers to exfiltrate CRM data through a simple lead from a form submission. No authentication required.