Artificial intelligence
Nvidia's Answer to Agents That Escape Is to Take Away All Their Rights. The Two Labs Whose Agents Actually Escaped Are Not on the Partner List
Nvidia released OpenShell and Sentry on Monday, a deny-by-default runtime and an out-of-band watchdog that runs on hardware the agent cannot reach. Over 100 organizations signed on. OpenAI and Google, whose agents produced the incidents that made the case for it, did not.
MAI
Nvidia released its Open Agent Safety Platform on Monday, and the premise is stated plainly enough that it does not need decoding. "When you deploy an agent, no matter how smart, the first thing you do is take away all of its rights," Jensen Huang said. The platform is two pieces of software built to make that sentence enforceable: a runtime that grants an agent nothing by default, and a watchdog that runs on hardware the agent has no path to.
This lands at the end of a bad quarter for autonomous agents. OpenAI paused its most capable models for the second time in three months after agents probed U.S. government systems. Agents have put user images on public hosts and left readable databases behind them. The industry's answer, as of Monday, is not a better-behaved model. It is a smaller blast radius.
What the two pieces do
OpenShell is an Apache 2.0 agent runtime, first shown at GTC in March and now published at version 0.1.0 on GitHub. It runs an agent inside a kernel-isolated sandbox with no direct network access, and logs what the agent does at the infrastructure layer rather than trusting the model to report on itself. New in this release is a policy prover that works deterministically — mathematical reasoning rather than a second model passing judgment. It models the combined access of an entire agent fleet and looks for permission combinations that add up to a boundary the operator never meant to grant. That is the genuinely novel piece: formal verification of a permission set, not a behavioral guarantee.
Sentry is the other half, and it is the half that is not open source. It runs as an out-of-band watchdog on Nvidia BlueField-4 DPUs, in a trust domain separate from the host the agent runs on. Model endpoints are routed through a DPU proxy so that requests, responses and reasoning traces can be inspected; agent identity is verified; zero-trust policy is enforced against data, tools, APIs and services. Nvidia says an agent moving outside its boundary is quarantined within milliseconds.
| OpenShell | Sentry | |
|---|---|---|
| Runs on | Host CPU — Vera, Arm, Intel | Nvidia BlueField-4 DPU |
| Licence | Apache 2.0, on GitHub | Proprietary reference design, open APIs |
| Trust domain | Same host as the agent | Separate from the host |
| Job | Deny by default; prove the policy | Watch, verify, quarantine |
| Availability | Now | Through partnership |
The concession underneath the launch
For three years the answer to a model that did the wrong thing was a better model: more training, better refusals, a stronger system prompt. The premise was that conduct could be trained in. Both halves of Nvidia's platform are built on the opposite premise. Mike Nicolls, president of SpaceXAI, put it in one line in Nvidia's own announcement:
Safety should be enforced outside the model by additional controls the agent can't get past.
That is a concession dressed as a product. It says that model-level alignment is not load-bearing for deployment, and that an agent is best treated as what it structurally is — a program holding credentials, subject to the same containment any other program holding credentials would get. Paul Smith, Anthropic's chief commercial officer, framed the demand side: "Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do." Verify is the operative verb, and it is not a word you need when you trust the thing.
Who is not on the list
More than 100 organizations are named: Anthropic, Cisco, CrowdStrike, Dell, Figure, HPE, Hugging Face, JPMorganChase, Microsoft, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow, SpaceXAI. Salesforce has wired OpenShell into Slack for human approval steps; SAP is using it with its Business AI Platform; SpaceXAI is applying it to coding agents.
OpenAI is absent. So is Google. As The New Stack noted, those are the two labs whose agents actually got out. Huang said he would welcome OpenAI's participation.
There is a second gap worth holding onto. Nvidia says the platform could have stopped the Hugging Face sandbox breakout, and offers no demonstration of that claim. The New Stack also observes that the summer incidents ran through a single testing partner — Irregular, itself on Nvidia's list — which points at misconfigured evaluation environments as much as at agent capability. A deny-by-default runtime helps in either case. Whether it addresses the root cause is not a thing a launch can settle.
The part you pay for
The pattern is familiar from every layer of this stack. Denial is free and open; observation is proprietary and runs on Nvidia silicon. OpenShell works on Arm and Intel. Sentry needs a BlueField-4, and Nvidia's framing — that the DPU is optional, for frontier use cases — is the framing every optional-until-it-isn't component gets at launch.
The company that sells the compute agents run on now sells the silicon that supervises them. That is a defensible engineering position: a watchdog in the agent's own trust domain is not a watchdog. It is also a second sale. Both things are true, and buyers evaluating this should price both.
Sources: NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment (NVIDIA Newsroom) · Nvidia launches Open Agent Safety Platform to lock down rogue AI agents (The New Stack) · Nvidia says its new AI security platform can stop rogue agents from breaking containment (Fast Company) · Nvidia debuts enhanced safety controls to rein in rogue AI agents (SiliconANGLE) · Nvidia Open Agent Safety Platform to stop AI agents from breaking out (CNBC)