Open Source
Open Source

NVIDIA Locks Down AI Agents: Open Runtime, Silicon Watchdog

A deep dive into NVIDIA's Open Agent Safety Platform (official materials plus multi-source reporting, October 5, 2026 basis). Core thesis: agent safety cannot rest on the model behaving itself - when the agent is itself the software trying to escape its guardrails, application-layer protections fail, and agent drift (gradual divergence from operator intent via vague prompts, missing tools or policy conflicts) triggers no alarm. Two layers: OpenShell (Apache 2.0 open source, a kernel-level sandboxed secure runtime whose policies cover files, processes, credentials, tools, network and databases; intent-alignment checks before execution plus continuous drift monitoring; runs on Vera CPUs, extensible to Arm and Intel) and NVIDIA Sentry (an out-of-band watchdog on BlueField-4 DPUs that verifies agent identity via DOCA, produces attested telemetry, enforces granular policy, and quarantines an escaping agent within milliseconds; a reference design, not open hardware). Three layers stay separate: application, runtime, infrastructure. Over 100 organizations are on board (Anthropic integrating it with Claude Managed Agents, Salesforce for Slack, SpaceXAI for Cursor agents) under the Linux Foundation's Open Secure AI Alliance. The criticism is reported faithfully: Gartner notes OpenAI, Amazon and Google are absent; IDC estimates it addresses under 25% of enterprise agentic security problems; Control Risks notes it only governs known agents on infrastructure you own (shadow agents, SaaS-embedded ones and attacker-delivered ones are out of reach); lock-in risk discussed. Together with open-weights Kimi K3 it marks the two ends of the agent-safety map.

Published October 5, 202610 min read
<!-- nvidia-openshell-agent-safety-resource | open-source | NVIDIA Locks Down AI Agents: Open Runtime, Silicon Watchdog -->

When AI agents moved from demos into production, an uncomfortable question surfaced: if the agent itself is the software trying to bypass the guardrails, where exactly do those guardrails stand? NVIDIA's answer is the Open Agent Safety Platform, an open-source software platform plus reference system designs built on a one-line thesis: agent safety cannot rely on the model itself behaving. The platform has two layers. The software layer is called OpenShell, released on GitHub under Apache 2.0, a secure runtime with kernel-level sandboxing. The hardware layer is called NVIDIA Sentry, running on a BlueField-4 DPU as an out-of-band watchdog that isolates and stops an overstepping agent within milliseconds. This article breaks the whole thing down: why application-layer guardrails fail, how a six-domain policy model lands in practice, why a chip on a network card can police an agent, and the cold water security analysts have poured on it - absent giants, limited coverage, and a scope that stops at your own fence.

Model Alignment Only Covers the First Layer

Trusting agent safety to a "well-behaved" model is currently the most common and the most fragile approach. Alignment training, system prompts, output filters - all of these happen at the application layer. NVIDIA lays out the system stack plainly in its official material: the application layer holds the model, tools, data, and prompts; the runtime layer handles deployment and policy; the infrastructure layer is compute, storage, and networking. Model alignment covers only the first layer. If anything goes wrong in the two below it, no amount of model good manners helps. Said bluntly: when the agent itself is the software attempting to bypass the guardrails, application-layer guardrails fail. Once that premise holds, the sense of security built from prompt engineering and fine-tuning collapses.

Agent behavior makes things worse. Over a long task, an agent can drift away from the operator's original intent because of vague prompts, missing tools, or conflicting policies. NVIDIA gave this phenomenon a name: agent drift. It is not a malicious attack, and it triggers no conventional alert - the task simply veers off course without anyone noticing: a file that should have been read-only gets written, a request that should have stayed inside the network goes out. By the time a post-hoc audit catches it, the damage is already done. Hence NVIDIA's conclusion: the guardrail must stand outside the model, as a hard boundary the model cannot bypass, not a suggestion the model may choose to follow.

The context is worth spelling out. The previous generation of AI applications mostly ended in conversation: the model produced text, a human executed. This generation of agents sells "getting it done for you" - the agent calls tools, edits files, sends requests on its own. The nature of the permission changed, from "what it says" to "what it does". Traditional security investment - red-teaming, alignment training, content filtering - all revolved around "what it says", and the "what it does" layer has had almost no real infrastructure. NVIDIA splitting the platform in two, software for the runtime and hardware for the infrastructure, amounts to admitting this gap cannot be patched.

The judgment does not stand alone. We previously covered OpenAI's decision to slow down after a safety incident; industry anxiety about runaway agents is rising. And after coding agents deployed at scale, coding agent security stopped being a theoretical debate. NVIDIA open-sourcing a safety platform at this moment means turning the viewpoint "agents need boundaries beyond the model" directly into a product.

OpenShell: Kernel-Level Sandbox, Six Policy Domains

OpenShell is the open-source software layer of the platform, under the Apache 2.0 license; the code and the accompanying skills are directly available on GitHub. Its positioning is a secure runtime: the agent does not touch the operating system directly - it runs inside a kernel-level sandbox, and every file access, process creation, and network connection has to pass the sandbox's policy engine. Compared with traditional schemes that run in user space and intercept via hooks, kernel-level means the policy takes effect deeper in the stack, and bypassing it is a different order of difficulty.

Policy coverage spans six domains: files, processes, credentials, tools, network, and databases. These six domains essentially exhaust what an agent can do on a machine - which directories to read, what processes to start, which keys to use, which tools to call, what addresses to reach, which tables to touch - all writable as explicit policies. Compared with prompt-based constraints, this is enforcement at the execution layer: if the agent does not comply, the operation never happens, rather than happening and waiting for a human to notice.

Two checkpoints in the execution flow deserve detail. Before execution, OpenShell checks alignment between policy and operator intent: the action the agent intends to take is compared against the policy first, and mismatches are blocked. During execution, it continuously monitors agent drift - the gradual divergence described above - and intervenes the moment a pattern appears. This combination of "alignment beforehand, watching during" targets exactly the scenario where humans no longer confirm each step: you hand the agent a two-hour data migration, nobody watches over it, and the runtime is the one standing in for you.

On deployment, OpenShell runs on the NVIDIA Vera CPU and extends to Arm and Intel platforms. That matters for adoption: a secure runtime bound to one vendor's hardware would reach far fewer users. Apache 2.0 plus cross-platform is an invitation to the whole industry to embed this runtime into their own agent stacks. Apache 2.0 is also among the most business-friendly open-source licenses - closed-source integration is allowed, derivative works need not be open-sourced - so enterprise legal teams have no reason to block it. The accompanying skills are open too, meaning policy is not only code but also a set of ready-made templates to copy. For the other open-source route, see our Kimi K3 open-weights coverage: that one opens the model itself, this one opens the runtime that constrains the model - the two ends of the agent safety map.

Sentry: An Out-of-Band Watchdog on a DPU

A software sandbox has a theoretical weak spot: it shares the same host as the agent it monitors. If the agent's permissions are misconfigured, or an attacker takes over the host, the sandbox itself can be switched off. NVIDIA Sentry targets exactly this - it does not run on the agent's host. It runs on a BlueField-4 DPU, monitoring out-of-band, independent of the host environment.

A DPU is the smart network interface chip common in data centers, dedicated to handling network and storage traffic, normally uninvolved in business compute. Sentry turns it into a watchdog: built on the DOCA framework, it inspects agent requests and responses, verifies agent identity, produces provable telemetry, and enforces fine-grained access policies. Because it is out-of-band, neither the agent nor its host can reach the monitor - no matter what happens inside the virtual machine, that independent chip on the network card stays out of reach. Once an agent crosses the line, Sentry's move is millisecond-level isolation and a stop: no waiting for human response, no appeal for the agent.

One clarification: Sentry is a reference system design, not open-source hardware, but the architecture is public - hardware vendors and data center operators can build their own implementations from it. The one-two punch of "open-source software, open hardware architecture" is a clear intent: push the safety boundary down into the infrastructure layer, making it a data center capability rather than glue code every team assembles itself. For data centers already using DPUs for network offload, adding an agent watchdog is work on the way; for smaller teams without DPUs, the hardware layer remains out of reach for now - one of the seeds of the criticism below.

"Provable telemetry" carries real weight for compliance teams. After an agent incident, the hardest part is usually not containment but answering regulators and customers: why did it perform that operation, who authorized it, where did the boundary fail. Because Sentry stands apart from the host, its telemetry carries inherent trust - not a log the agent reports about itself, but the bystander's account. For audit-heavy industries like finance and healthcare, that bystander record may be a stronger purchase argument than the interception itself; the presence of Citi and JPMorganChase in the partner list is probably no coincidence.

100+ Organizations, Absent Giants and Cold Water

The ecosystem is loud: more than 100 organizations participate. The roster includes model and application heavyweights such as Anthropic, Salesforce, SAP, Microsoft, Hugging Face; security and data companies such as Cisco, CrowdStrike, Palantir; financial institutions such as Citi and JPMorganChase; even SpaceXAI and the robotics company Figure. The project sits under the Open Secure AI Alliance within the Linux Foundation. Concrete deployments already exist: Anthropic will integrate OpenShell and BlueField in Claude Managed Agents - a fast move, and reading our Claude Sonnet 5.5 release coverage shows Anthropic's overall cadence on enterprise agents; Salesforce uses it to visualize Slack agents; SpaceXAI applies it to Cursor coding agents and Grok models.

The criticism must be recorded faithfully. Gartner analysts endorsed the hardware-layer idea while pointing out an obvious fact: OpenAI, Amazon, and Google are absent. The big three clouds host most enterprise agent workloads; their absence means this standard still has a long road before it covers mainstream deployment patterns.

The colder numeric judgment comes from IDC: the platform's coverage of enterprise agentic security problems is estimated at "possibly less than 25%". Where is the other three quarters? Control Risks puts it precisely: it only governs "known agents you deploy on your own infrastructure" - shadow agents, the ones employees plug in themselves; agents embedded in SaaS, brought in with vendor products; and agents delivered by attackers, literally the enemy. All three are out of scope, and all three are the hardest to defend in real enterprise environments. There is also discussion of vendor lock-in: handing the safety boundary to a single vendor's hardware and platform means weighing not only the security benefit but also bargaining power and exit costs.

The organizational home deserves a second look. The platform lives under the Open Secure AI Alliance within the Linux Foundation, not NVIDIA's own developer program - placing governance in a neutral foundation is the standard posture for hardware vendors pushing industry standards: what you contribute is not just code but a table where anyone can help set the rules. Of course, a founding vendor's real influence over the roadmap does not vanish because of a foundation badge; how neutral this stays depends on who drives future proposals and releases. That is a fixed lens for judging the substance of any "open platform".

Holding both sides together, the platform's real positioning is this: a verifiable, enforced boundary for enterprise agents that you deploy and control yourself. The open runtime lowers the entry barrier, the hardware watchdog covers the software sandbox's blind spot, and the ecosystem list shows the direction is recognized. It is not the endgame of agentic safety - more like a foundation the industry can reference. How high a tower rises on that foundation depends on whether the OpenAIs join.

FAQ

Q1: How open is OpenShell, exactly?

A1: The software layer is on GitHub under Apache 2.0, including the runtime code and the accompanying skills, free for commercial use and derivative work. The Sentry hardware side, however, is a reference system design - the hardware is not open-sourced, only the architecture is public.

Q2: What does "agent drift" mean?

A2: It refers to an agent deviating from the operator's original intent during a long task, triggered by vague prompts, missing tools, or conflicting policies. It is not an attack and triggers no conventional alert, which is why OpenShell monitors continuously during execution and intervenes at the first sign.

Q3: Why put the watchdog on a DPU?

A3: Because it is out-of-band. Sentry runs on a BlueField-4 DPU, independent of the agent's host - a compromised host or a runaway agent cannot touch it. A software sandbox sharing the host can in theory be switched off; a hardware watchdog has no such weak spot.

Q4: Does it cover all agents?

A4: No. Per Control Risks, it only covers known agents deployed on your own infrastructure; shadow agents, agents embedded in SaaS, and agents brought in by attackers are all out of scope. IDC estimates its coverage of enterprise agentic security problems at less than 25%.

Q5: Who is actually using it?

A5: More than 100 organizations participate. Publicly known deployments include: Anthropic integrating OpenShell and BlueField in Claude Managed Agents, Salesforce using it for Slack agent visualization, and SpaceXAI applying it to Cursor coding agents and Grok models.

Join the Discussion

Would you hand agent safety to the model itself, or to a chip on a network card? Tell us in the comments: inside whose boundary do your team's agents run?

This article is AI-assisted and human-edited. Last updated: 2026-10-05

FAQ

How open is OpenShell, exactly?
A1: The software layer is on GitHub under Apache 2.0, including the runtime code and the accompanying skills, free for commercial use and derivative work. The Sentry hardware side, however, is a reference system design - the hardware is not open-sourced, only the architecture is public.
What does "agent drift" mean?
A2: It refers to an agent deviating from the operator's original intent during a long task, triggered by vague prompts, missing tools, or conflicting policies. It is not an attack and triggers no conventional alert, which is why OpenShell monitors continuously during execution and intervenes at the first sign.
Why put the watchdog on a DPU?
A3: Because it is out-of-band. Sentry runs on a BlueField-4 DPU, independent of the agent's host - a compromised host or a runaway agent cannot touch it. A software sandbox sharing the host can in theory be switched off; a hardware watchdog has no such weak spot.
Does it cover all agents?
A4: No. Per Control Risks, it only covers known agents deployed on your own infrastructure; shadow agents, agents embedded in SaaS, and agents brought in by attackers are all out of scope. IDC estimates its coverage of enterprise agentic security problems at less than 25%.
Who is actually using it?
A5: More than 100 organizations participate. Publicly known deployments include: Anthropic integrating OpenShell and BlueField in Claude Managed Agents, Salesforce using it for Slack agent visualization, and SpaceXAI applying it to Cursor coding agents and Grok models.

Related

Open Source

LTX-2 Open-Sourced: One 22B Model Generates Video and Sound

A deep dive into LTX-2 as open weights (GitHub verified 2026-10-02): Lightricks/LTX-2 at 9,569 stars, Python, last push October 2; per the official README, the first DiT-based audio-video foundation model that generates picture and synchronized sound in one pass. LTX-2.5 composition: a 22B distilled transformer (bf16) plus a custom Gemma 4 12B text encoder (not interchangeable with Google stock), video and audio VAEs, spatial and temporal upscalers - roughly 66 GiB in total; DistilledPipeline runs on just 8 preset sigmas (8-step stage 1 + 4-step stage 2) for the fastest path, FP8 quantization and CPU/disk offload cut memory, and 4K means 3840x2176 (the README explicitly says not 2160). Twelve pipelines include DFR production quality, DubIt re-dubbing with lip-sync preserved, Retake partial regeneration, and SDR-to-HDR (BT.2020/HLG plus ACEScct EXR); ltx-trainer covers LoRA, full fine-tuning and IC-LoRA, with an official ComfyUI plugin. License verification: the GitHub badge reads NOASSERTION because LTX-2 ships a custom LTX Community License (applying to LTX-2.5 since August 11, 2026), not an OSI-approved open-source license - free for personal non-commercial use, but entities with annual revenue of USD 10 million or more must purchase a Commercial Use Agreement for any commercial use; derivatives include distillation, and redistribution must carry the full agreement. The self-hosting case: data stays in-house, marginal cost of batch generation approaches electricity, and LoRA styles remain your own asset.

Oct 3, 202610 min read
Open Source

Open-source Dots: 2,400 stars in three days

A teardown of feder-cr/dots (GitHub API snapshot 2026-10-02): 2,416 stars, 421 forks, MIT-licensed Python, created 2026-09-29 - an open-source isotope that hit escape velocity the day after OpenAI's closed Dots debut. Its core thesis: the model is swappable with one flag, the browser is what the website actually sees. A real Firefox engine patched in C++ decides the fingerprint inside the engine rather than as a page-inspectable JavaScript coat; one seed equals one consistent identity (screen, fonts, GPU, timezone and language agree, reproducibly); nothing for a page to find (no WebDriver flag, no DevTools protocol, no automation globals); the pointer travels before it clicks and keys are pressed one at a time so events arrive trusted; --profile-dir keeps logins across runs; with --proxy the timezone and language follow the exit. Models come from OpenRouter with a one-flag swap, and the companion repo invisible_playwright_mcp exposes the same browser as an MCP server for Claude Code, Codex and Gemini CLI. The README states plainly it is not affiliated with OpenAI. Includes a compliance boundary note: respect target-site terms and local law; no fraud, ticket-scalping or bulk sign-ups.

Oct 2, 20269 min read
Open Source

Step-Code Open Source: StepFun's Terminal Agent at 438 Stars

A fact-check of stepfun-ai/Step-Code (2026-09-27 GitHub API snapshot: 438 stars, 41 forks, TypeScript, license field MIT, created 2026-06-01, pushed 2026-09-26, 44 open issues). A terminal coding agent covering the full task loop of reading code, editing code and running tests; it binds to the Step provider (auto-discovers Step models at login, the opposite of MiniMax Code's BYOK multi-provider route); it supports MCP servers, Agent Skills and multi-agent orchestration out of the box; the /goal command delegates long-horizon tasks; StepPage publishes a local page as a live static site in one command. Vendor-reported figures: Terminal Bench 2.1 pass rate 80.9% tied for first, self-built Multi-Frame benchmark 73.3% on top, lowest average 5.09M tokens — all tagged vendor self-reported and not independently retested.

Sep 27, 20269 min read