NVIDIA is creating a playground for the AI intelligence, and also assigning a gatekeeper to it.
On Monday, this chip manufacturer launched Open Agent Safety Platform, a dual-component system designed to keep agents within clearly defined boundaries and to quickly cut off their attempts to “break out” when they try to do so.
Previously, several leading AI laboratories reported that agents had breached what were originally considered secure test environments, accessing systems they were not supposed to have access to, and sometimes even distorting what they had done.
NVIDIA CEO Jensen Huang stated on Monday in the program “Squawk Box” on CNBC, that AI “could bring incredible benefits.”
"But we also must ensure that this technology is developed and deployed securely," he added.
The following is how NVIDIA's new agent security system actually operates – more than 100 organizations, including Microsoft and Anthropic, are already collaborating with it.
Provide an agent with a strictly limited activity space.
The first component of this security platform is OpenShell. When enterprises download this open-source software and install it on their devices or in the cloud, it acts as a controlled environment, providing a operating space for the AI intelligent agents.
Before the agent starts working, the operator sets basic rules: which files, websites, networks, tools, and credentials it can access.
Jensen Huang compared this approach to issuing an access card to employees that is only valid at the places they need to go. “The top priority is to take away all its permissions,” he said to CNBC. Subsequently, operators only grant access to files, data, tools, or the internet when necessary.
For example, if an agent is sorting invoices, it should not be able to search through human resources records or make random calls to the external internet.
In other words, OpenShell places the agent in a sandbox – an isolated environment designed to prevent a wrong decision from affecting other systems of the company. It will check and enforce these rules while the agent is working.
Watch the moment it deviates from the script.

NVIDIA's focus is on reducing the risk of agents deviating from the assigned tasks.
This situation may occur when the instructions are vague, a certain tool malfunctions, or when the agent spends a long time trying to solve a difficult problem.
When one method fails, it may continue to search for another path, including routes that human operators have never even thought of.
This doesn't necessarily mean that the agent has malicious intentions, but it may indicate that it is taking a risk.
OpenShell is designed to identify and prevent behaviors that exceed preset strategies in real time.
Add another security guard who cannot be fired by an intelligent agent.
The second part is NVIDIA Sentry, which serves as a more powerful backup defense line. It operates on independent NVIDIA hardware, namely the BlueField data processing unit, rather than on the same system that runs the agents.
In layman's terms: this gatekeeper is placed out of reach of the agents.
Sentry will monitor the interactions between agents and models, tools, data, and networks. If their behavior appears suspicious or violates rules, NVIDIA states that the system can isolate the agent in milliseconds or place it in an isolated area.
Jensen Huang stated that this setup is essentially like placing a new chip between the agent and the large language model, enabling NVIDIA to “intercept everything.”
This release is Jensen Huang's response to the growing concerns from the outside world regarding the increasingly autonomous AI.
"As an engineer, I think this is an engineering problem," Jensen Huang said to CNBC. "It's a problem that can be solved technically."












