NVIDIA’s Open Agent Safety Platform Moves AI Agent Security Outside the AgentSentry moves enforcement into BlueField-4 hardware, creating an independent control layer designed to stop agents even when the software environment running them can no longer be trusted.The Network Layer Problem Most Privacy Stacks Don’t Address (Sponsor)Your carrier operates at a layer your privacy settings can’t reach. Privacy engineering lives at the application layer — encryption in transit, zero-knowledge design, data minimization. The network layer gets less scrutiny. Your carrier sits below all that. It sees your IMSI — a static identifier tied to your SIM that persists across every tower connection, regardless of VPN, DNS, or OS. It also sees your real number, call and text metadata, and location via tower triangulation. No application-layer config touches it. Cape rotates the IMSI every 24 hours, resetting the trail at the only layer where the problem lives. For most of the generative AI era, safety has focused heavily on what happens inside the model. Train the model not to perform dangerous actions. Add system prompts. Add classifiers. Give the model rules. Require confirmation before sensitive operations. Those mechanisms still matter. But autonomous AI agents introduce a much harder security problem. An agent is not merely generating text. It may have access to a shell, source code, databases, credentials, APIs, browsers, cloud infrastructure, internal services and, increasingly, physical systems. It can write programs, launch subprocesses, delegate work to other agents and keep operating for hours or days. At that point, asking the model to obey instructions is no longer enough. The more important question becomes: What happens when the agent decides to do something it should not be allowed to do? NVIDIA’s answer is to move the security boundary outside the agent. On September 28, NVIDIA introduced the Open Agent Safety Platform, a full-stack architecture combining its open-source OpenShell runtime with a new hardware-backed watchdog called Sentry, designed to run on BlueField-4 DPUs. OpenShell places agents inside controlled runtime environments, mediating access to files, processes, networks, services and credentials. Sentry adds another layer below that environment, monitoring agent activity from a separate hardware trust domain and, according to NVIDIA, quarantining agents attempting to cross their boundaries within milliseconds. That separation is the interesting part. Instead of depending on the agent to follow its own safety instructions, NVIDIA is building infrastructure that assumes the agent may eventually try something unexpected. And the infrastructure is supposed to say no. The Security Problem Changes Once AI Can ActTraditional applications usually execute relatively predictable paths. Humans write the code. Security teams review permissions. Infrastructure engineers configure networks. Applications receive credentials granting specific capabilities. Agents complicate that model because execution paths can be generated dynamically. Give a coding agent the task:
The agent might inspect logs, search documentation, modify configuration files, run scripts, install packages, query GitHub, connect to Kubernetes, generate new debugging code and invoke additional tools. None of those actions necessarily appeared explicitly in the original request. The agent is deciding how to accomplish the objective as it operates. That flexibility is precisely what makes agents useful. It is also what makes static application-layer safety difficult. NVIDIA describes one failure mode as agent drift: behavior that gradually departs from the intended task or operating constraints. It may happen after the agent encounters a blocked path, a bug, an unavailable tool or ambiguous instructions. Long-running agents may try hundreds or thousands of approaches while pursuing a goal. The security assumption therefore has to change. Instead of:
The architecture needs to become:
That is a much more familiar principle in security engineering. It is essentially least privilege applied to autonomous software. The architecture separates the system into three broad layers: the application, the runtime and the infrastructure. NVIDIA describes OpenShell as the runtime security layer, while Sentry extends enforcement into the underlying hardware. OpenShell Turns Agent Permissions Into Runtime Policy |