When AI agents moved from demos into production, an uncomfortable question surfaced: if the agent itself is the software trying to bypass the guardrails, where exactly do those guardrails stand? NVIDIA's answer is the Open Agent Safety Platform, an open-source software platform plus reference system designs built on a one-line thesis: agent safety cannot rely on the model itself behaving. The platform has two layers. The software layer is called OpenShell, released on GitHub under Apache 2.0, a secure runtime with kernel-level sandboxing. The hardware layer is called NVIDIA Sentry, running on a BlueField-4 DPU as an out-of-band watchdog that isolates and stops an overstepping agent within milliseconds. This article breaks the whole thing down: why application-layer guardrails fail, how a six-domain policy model lands in practice, why a chip on a network card can police an agent, and the cold water security analysts have poured on it - absent giants, limited coverage, and a scope that stops at your own fence.
Model Alignment Only Covers the First Layer
Trusting agent safety to a "well-behaved" model is currently the most common and the most fragile approach. Alignment training, system prompts, output filters - all of these happen at the application layer. NVIDIA lays out the system stack plainly in its official material: the application layer holds the model, tools, data, and prompts; the runtime layer handles deployment and policy; the infrastructure layer is compute, storage, and networking. Model alignment covers only the first layer. If anything goes wrong in the two below it, no amount of model good manners helps. Said bluntly: when the agent itself is the software attempting to bypass the guardrails, application-layer guardrails fail. Once that premise holds, the sense of security built from prompt engineering and fine-tuning collapses.
Agent behavior makes things worse. Over a long task, an agent can drift away from the operator's original intent because of vague prompts, missing tools, or conflicting policies. NVIDIA gave this phenomenon a name: agent drift. It is not a malicious attack, and it triggers no conventional alert - the task simply veers off course without anyone noticing: a file that should have been read-only gets written, a request that should have stayed inside the network goes out. By the time a post-hoc audit catches it, the damage is already done. Hence NVIDIA's conclusion: the guardrail must stand outside the model, as a hard boundary the model cannot bypass, not a suggestion the model may choose to follow.
The context is worth spelling out. The previous generation of AI applications mostly ended in conversation: the model produced text, a human executed. This generation of agents sells "getting it done for you" - the agent calls tools, edits files, sends requests on its own. The nature of the permission changed, from "what it says" to "what it does". Traditional security investment - red-teaming, alignment training, content filtering - all revolved around "what it says", and the "what it does" layer has had almost no real infrastructure. NVIDIA splitting the platform in two, software for the runtime and hardware for the infrastructure, amounts to admitting this gap cannot be patched.
The judgment does not stand alone. We previously covered OpenAI's decision to slow down after a safety incident; industry anxiety about runaway agents is rising. And after coding agents deployed at scale, coding agent security stopped being a theoretical debate. NVIDIA open-sourcing a safety platform at this moment means turning the viewpoint "agents need boundaries beyond the model" directly into a product.
OpenShell: Kernel-Level Sandbox, Six Policy Domains
OpenShell is the open-source software layer of the platform, under the Apache 2.0 license; the code and the accompanying skills are directly available on GitHub. Its positioning is a secure runtime: the agent does not touch the operating system directly - it runs inside a kernel-level sandbox, and every file access, process creation, and network connection has to pass the sandbox's policy engine. Compared with traditional schemes that run in user space and intercept via hooks, kernel-level means the policy takes effect deeper in the stack, and bypassing it is a different order of difficulty.
Policy coverage spans six domains: files, processes, credentials, tools, network, and databases. These six domains essentially exhaust what an agent can do on a machine - which directories to read, what processes to start, which keys to use, which tools to call, what addresses to reach, which tables to touch - all writable as explicit policies. Compared with prompt-based constraints, this is enforcement at the execution layer: if the agent does not comply, the operation never happens, rather than happening and waiting for a human to notice.
Two checkpoints in the execution flow deserve detail. Before execution, OpenShell checks alignment between policy and operator intent: the action the agent intends to take is compared against the policy first, and mismatches are blocked. During execution, it continuously monitors agent drift - the gradual divergence described above - and intervenes the moment a pattern appears. This combination of "alignment beforehand, watching during" targets exactly the scenario where humans no longer confirm each step: you hand the agent a two-hour data migration, nobody watches over it, and the runtime is the one standing in for you.
On deployment, OpenShell runs on the NVIDIA Vera CPU and extends to Arm and Intel platforms. That matters for adoption: a secure runtime bound to one vendor's hardware would reach far fewer users. Apache 2.0 plus cross-platform is an invitation to the whole industry to embed this runtime into their own agent stacks. Apache 2.0 is also among the most business-friendly open-source licenses - closed-source integration is allowed, derivative works need not be open-sourced - so enterprise legal teams have no reason to block it. The accompanying skills are open too, meaning policy is not only code but also a set of ready-made templates to copy. For the other open-source route, see our Kimi K3 open-weights coverage: that one opens the model itself, this one opens the runtime that constrains the model - the two ends of the agent safety map.
Sentry: An Out-of-Band Watchdog on a DPU
A software sandbox has a theoretical weak spot: it shares the same host as the agent it monitors. If the agent's permissions are misconfigured, or an attacker takes over the host, the sandbox itself can be switched off. NVIDIA Sentry targets exactly this - it does not run on the agent's host. It runs on a BlueField-4 DPU, monitoring out-of-band, independent of the host environment.
A DPU is the smart network interface chip common in data centers, dedicated to handling network and storage traffic, normally uninvolved in business compute. Sentry turns it into a watchdog: built on the DOCA framework, it inspects agent requests and responses, verifies agent identity, produces provable telemetry, and enforces fine-grained access policies. Because it is out-of-band, neither the agent nor its host can reach the monitor - no matter what happens inside the virtual machine, that independent chip on the network card stays out of reach. Once an agent crosses the line, Sentry's move is millisecond-level isolation and a stop: no waiting for human response, no appeal for the agent.
One clarification: Sentry is a reference system design, not open-source hardware, but the architecture is public - hardware vendors and data center operators can build their own implementations from it. The one-two punch of "open-source software, open hardware architecture" is a clear intent: push the safety boundary down into the infrastructure layer, making it a data center capability rather than glue code every team assembles itself. For data centers already using DPUs for network offload, adding an agent watchdog is work on the way; for smaller teams without DPUs, the hardware layer remains out of reach for now - one of the seeds of the criticism below.
"Provable telemetry" carries real weight for compliance teams. After an agent incident, the hardest part is usually not containment but answering regulators and customers: why did it perform that operation, who authorized it, where did the boundary fail. Because Sentry stands apart from the host, its telemetry carries inherent trust - not a log the agent reports about itself, but the bystander's account. For audit-heavy industries like finance and healthcare, that bystander record may be a stronger purchase argument than the interception itself; the presence of Citi and JPMorganChase in the partner list is probably no coincidence.
100+ Organizations, Absent Giants and Cold Water
The ecosystem is loud: more than 100 organizations participate. The roster includes model and application heavyweights such as Anthropic, Salesforce, SAP, Microsoft, Hugging Face; security and data companies such as Cisco, CrowdStrike, Palantir; financial institutions such as Citi and JPMorganChase; even SpaceXAI and the robotics company Figure. The project sits under the Open Secure AI Alliance within the Linux Foundation. Concrete deployments already exist: Anthropic will integrate OpenShell and BlueField in Claude Managed Agents - a fast move, and reading our Claude Sonnet 5.5 release coverage shows Anthropic's overall cadence on enterprise agents; Salesforce uses it to visualize Slack agents; SpaceXAI applies it to Cursor coding agents and Grok models.
The criticism must be recorded faithfully. Gartner analysts endorsed the hardware-layer idea while pointing out an obvious fact: OpenAI, Amazon, and Google are absent. The big three clouds host most enterprise agent workloads; their absence means this standard still has a long road before it covers mainstream deployment patterns.
The colder numeric judgment comes from IDC: the platform's coverage of enterprise agentic security problems is estimated at "possibly less than 25%". Where is the other three quarters? Control Risks puts it precisely: it only governs "known agents you deploy on your own infrastructure" - shadow agents, the ones employees plug in themselves; agents embedded in SaaS, brought in with vendor products; and agents delivered by attackers, literally the enemy. All three are out of scope, and all three are the hardest to defend in real enterprise environments. There is also discussion of vendor lock-in: handing the safety boundary to a single vendor's hardware and platform means weighing not only the security benefit but also bargaining power and exit costs.
The organizational home deserves a second look. The platform lives under the Open Secure AI Alliance within the Linux Foundation, not NVIDIA's own developer program - placing governance in a neutral foundation is the standard posture for hardware vendors pushing industry standards: what you contribute is not just code but a table where anyone can help set the rules. Of course, a founding vendor's real influence over the roadmap does not vanish because of a foundation badge; how neutral this stays depends on who drives future proposals and releases. That is a fixed lens for judging the substance of any "open platform".
Holding both sides together, the platform's real positioning is this: a verifiable, enforced boundary for enterprise agents that you deploy and control yourself. The open runtime lowers the entry barrier, the hardware watchdog covers the software sandbox's blind spot, and the ecosystem list shows the direction is recognized. It is not the endgame of agentic safety - more like a foundation the industry can reference. How high a tower rises on that foundation depends on whether the OpenAIs join.
FAQ
Q1: How open is OpenShell, exactly?
A1: The software layer is on GitHub under Apache 2.0, including the runtime code and the accompanying skills, free for commercial use and derivative work. The Sentry hardware side, however, is a reference system design - the hardware is not open-sourced, only the architecture is public.
Q2: What does "agent drift" mean?
A2: It refers to an agent deviating from the operator's original intent during a long task, triggered by vague prompts, missing tools, or conflicting policies. It is not an attack and triggers no conventional alert, which is why OpenShell monitors continuously during execution and intervenes at the first sign.
Q3: Why put the watchdog on a DPU?
A3: Because it is out-of-band. Sentry runs on a BlueField-4 DPU, independent of the agent's host - a compromised host or a runaway agent cannot touch it. A software sandbox sharing the host can in theory be switched off; a hardware watchdog has no such weak spot.
Q4: Does it cover all agents?
A4: No. Per Control Risks, it only covers known agents deployed on your own infrastructure; shadow agents, agents embedded in SaaS, and agents brought in by attackers are all out of scope. IDC estimates its coverage of enterprise agentic security problems at less than 25%.
Q5: Who is actually using it?
A5: More than 100 organizations participate. Publicly known deployments include: Anthropic integrating OpenShell and BlueField in Claude Managed Agents, Salesforce using it for Slack agent visualization, and SpaceXAI applying it to Cursor coding agents and Grok models.
Join the Discussion
Would you hand agent safety to the model itself, or to a chip on a network card? Tell us in the comments: inside whose boundary do your team's agents run?