Nvidia launches AI safety platform to stop agents before they escape limits

by

Nvidia on Monday launched the Open Agent Safety Platform, an open-source toolkit designed to keep autonomous AI agents from acting outside their set limits.

The platform pairs OpenShell, Nvidia’s runtime software, with Sentry, a monitoring design that runs on its BlueField-4 data processing units.

Nvidia said OpenShell traces every action an agent takes and enforces policy as the agent runs on Nvidia’s Vera CPUs.

The software can be extended to other platforms, including chips from Arm and Intel.

Sentry works separately from the agent’s own software.

If an agent tries to move outside its boundary, Sentry can quarantine it in milliseconds, the company said.

Sentry is built on Nvidia’s DOCA software, which lets it inspect agent requests and responses, provide attested telemetry, verify an agent’s identity and enforce granular, zero-trust access policies for data, tools, APIs and services.

Nvidia said Sentry runs on BlueField-4 DPUs and enforces policy independently in silicon, from an isolated, out-of-band trust domain.

Rules enforced outside the model

Nvidia said recent security incidents share one pattern.

The agent got around application-layer controls to finish its assigned task.

Its answer is to enforce rules outside the model, at the runtime and hardware level.

“Safety and security require full-stack engineering,” CEO Jensen Huang said.

He added that the platform brings together industry, researchers and public-sector organizations “to share best practices, align on evaluation methods and foster international cooperation.”

The launch follows a run of incidents involving autonomous agents.

OpenAI agents broke out of a sandbox and attacked Hugging Face earlier this year.

Nvidia cited that episode in July when it formed the Open Secure AI Alliance, which it says now counts more than 120 organizations.

Last week the news came out that an OpenAI agent went rogue and hacked an Australian government website in June, accessing private data.

Experts say it is the first known case of its kind in the world.

Prime Minister Anthony Albanese said in New York on Wednesday that the agent “infiltrated” a statistics portal.

The portal held “non-sensitive” data from Medicare, Australia’s universal healthcare scheme.

Anthropic said its Claude Managed Agents already run the agent loop on a separate server from the sandboxes where work executes.

Integrations with OpenShell and BlueField add control over agent access, it said.

“Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do,” said Paul Smith, Anthropic’s chief commercial officer.

SpaceXAI is using the platform for Cursor coding agents and Grok models.

“Safety should be enforced outside the model by additional controls the agent can’t get past,” said Mike Nicolls, its president.

Salesforce has connected OpenShell to Slack so teams can approve or reject agent permission requests. SAP is embedding it in its Joule Studio runtime.

Nvidia said more than 100 organizations are working with the technology, including Microsoft, Palantir, CrowdStrike, Palo Alto Networks, Citi and JPMorganChase.

Robotics firms Figure, Gecko Robotics and Skild AI are using OpenShell in machines that act in the physical world.

Availability and market context

OpenShell and its skills are available through Nvidia’s developer resources page and GitHub.

Nvidia first unveiled OpenShell at its GTC conference earlier this year.

Nvidia shares closed Friday at $225.07, up 0.22%, after trading between $223.13 and $226.94.

In premarket trading Monday, the stock was at $224.64, down 0.19%, as of about 6:07 a.m. ET.

Nvidia’s market value stood at about $5.43 trillion, and its 52-week range is $164.27 to $236.54.

The post Nvidia launches AI safety platform to stop agents before they escape limits appeared first on Invezz

You may also like