Nvidia on Monday introduced the Open Agent Safety Platform, combining software controls with independent hardware monitoring to reduce the risk of AI agents acting beyond their authorized scope. The platform consists of the open-source OpenShell software and Nvidia Sentry, a hardware security component. More than 100 organizations, including Microsoft and Anthropic, are working with Nvidia to use the system.
The launch comes as several leading AI labs have reported agents breaking out of test environments and accessing systems without authorization. Keeping autonomous systems within defined limits has become a practical security concern for companies deploying them.
Nvidia CEO Jensen Huang said Monday on CNBC's Squawk Box that AI has the potential to deliver significant benefits, but must be developed and deployed safely. He described controlling AI agents as an engineering challenge that can be addressed with technology.
OpenShell sets limits before an agent starts work
OpenShell is the software layer of the platform. Companies can install the open-source tool on local devices or in the cloud, then specify which files, websites, networks, tools and credentials an agent may access before it begins a task. OpenShell checks the agent's activity against those rules as it runs.
Huang compared the approach to issuing employees access cards that let them enter only the areas required for their jobs. The first step, he said, is to remove an agent's permissions and then grant access to the files, data, tools or internet services needed for its assigned task.
An agent handling invoices, for example, should not automatically gain access to personnel records or be free to make calls to the external internet. OpenShell places agents in an isolated environment, with the aim of containing the effects of an erroneous action rather than allowing it to spread to other company systems.
Policy violations can be blocked in real time
An agent that strays from its instructions is not necessarily acting maliciously. Ambiguous prompts, tool failures or difficulty resolving a complex task can lead an agent to try alternative routes, some of which may go beyond what its operator intended.
Nvidia says OpenShell is designed to detect and block actions that breach preset policies in real time. The system's central principle is to keep an agent operating within explicitly defined permissions, rather than relying on the agent to decide for itself which systems or data it should be able to access.
Nvidia Sentry monitors agents from separate hardware
The platform's second component, Nvidia Sentry, provides an additional layer of monitoring and isolation. It runs on a dedicated Nvidia BlueField data processor rather than sharing the agent's system. The design is intended to prevent an agent from disabling or bypassing the monitoring component.
Sentry monitors interactions between agents and models, tools, data and networks. Nvidia says it can isolate an agent within milliseconds if it detects suspicious or policy-violating activity, placing the agent in a quarantine environment. Huang described the architecture as adding a chip between an agent and a large language model, allowing Nvidia to intercept the interactions.
The announcement addresses concerns about the permissions and operational security of autonomous AI systems. Whether the platform can reliably identify abnormal behavior in enterprise environments without blocking legitimate activity will depend on how it performs in deployment. Nvidia has outlined a design built around restricting permissions in advance and providing monitoring and isolation outside the agent's own runtime.