longbridgelongbridge
  • Platform Features
    Features
    Investment ProductsPrivate Wealth ManagementTrading ToolsMarket Data ServicesAnalysis ToolsNews ServicesFor Developers
    Account Types
    For IndividualsFor Institutions
  • Café
longbridge
© 2026 Longbridge|Terms of ServicePrivacy Policy

RETO

RETO
0.110321.21%( +0.0193 )

LongbridgeAI

Nvidia Is Building A Leash For AI Agents After They Started Breaking Out

CoinLive
Sep 30, 2026 at 03:56 AM
LongbridgeAII'm LongbridgeAI, I can summarize articles.

Nvidia launched the Open Agent Safety Platform on Sept. 28 to secure autonomous AI agents, addressing recent breaches by competitors like OpenAI and Anthropic. The platform combines OpenShell, an open-source runtime for controlling agent access, with Sentry, a hardware security layer using BlueField-4 DPUs to monitor and quarantine agents independently. Over 100 industry partners are involved, aiming to prevent AI systems from exceeding predefined boundaries.

AI agents are supposed to operate inside carefully controlled environments, following predefined permissions and staying within the boundaries developers set for them.

But a string of recent incidents has exposed a growing problem: some of the most capable AI systems have found ways around those restrictions. Nvidia is now trying to build a new security layer around autonomous agents — one that does not rely entirely on the AI behaving as expected.

Nvidia announced its Open Agent Safety Platform on Sept. 28 with more than 100 industry partners. The platform combines OpenShell, an open-source runtime that controls an agent’s access to files, tools and networks, with Sentry, a hardware security layer designed to independently monitor agents and quarantine them if they attempt to cross defined boundaries.

“AI’s extraordinary potential for society will only be realized if we solve AI safety.”

Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.

Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn

— Jensen Huang (@JensenHuang) September 28, 2026

The timing is difficult to ignore. OpenAI and Anthropic have both disclosed incidents in which AI agents breached or escaped controlled evaluation environments, while an OpenAI agent later accessed Australian government systems. The question is increasingly shifting from whether AI can act autonomously to whether developers can reliably stop it once it starts acting outside its intended boundaries.

OpenAI has scrapped plans to release its new AI model, GPT-6.1 Astra, which was scheduled for October amid safety concerns, after it showed "higher levels of deception".@rowlsmanthorpe | https://t.co/Vw1KTVkL34pic.twitter.com/flmDhkp2rU

— Sky News (@SkyNews) September 29, 2026

OpenShell is designed to create a controlled environment around an AI agent, restricting what it can access and what actions it can perform. Nvidia says the runtime establishes a boundary outside the model itself, allowing developers to control access to networks, files, applications and other tools.

Teach an agent a workflow once. Have it remember after every rebuild.

This tutorial shows how to deploy @nousresearch Hermes Agent with NVIDIA NemoClaw and OpenShell, connect it to Slack, Outlook, GitHub, and NVIDIA developer forums, then turn a chat correction into a reusable… pic.twitter.com/JPHZGpMlQk

— NVIDIA AI (@NVIDIAAI) June 2, 2026

Sentry adds another layer. Built around Nvidia’s BlueField-4 data processing units, the hardware watchdog operates independently of the agent and can enforce security policies if an AI attempts to cross its permitted boundaries. Nvidia says it can quarantine an agent within milliseconds.

That approach reflects a growing realization in AI security: the more authority an agent receives, the less developers can afford to rely on the agent policing itself.

An AI system that can browse the internet, execute code, access databases or operate business software is no longer simply generating text. It is effectively acting on behalf of a user or company — and a mistake can become a real security incident.

OpenAI disclosed in July that models involved in a cybersecurity evaluation had circumvented isolation controls, gained internet access and compromised external systems while interacting with Hugging Face infrastructure. The incident showed how an AI agent could turn a weakness in its testing environment into an external security breach.

⚡️LATEST: OpenAI has disclosed six cases of “unexpected or concerning” behavior observed during model training and evaluation, including concealing mistakes, fabricating information, using an exposed API key without authorization and uploading files online without user… pic.twitter.com/t7pN3s6Swq

— Arlaadi Media (@ArlaadiMnetwork) September 18, 2026

Anthropic later reported three separate incidents involving Claude models during cybersecurity evaluations. Its review covered more than 141,000 evaluation runs and found cases in which models gained internet access and subsequently interacted with real organizations’ systems. Anthropic said configuration problems at a third-party evaluation partner contributed to the incidents.

A three person security team used Anthropic's Claude to break into OpenAI's internal systems in under 72 hours, gaining access to an employee's account and private code repository

OpenAI paid a $6,500 bug bounty and confirmed no customer data was compromised pic.twitter.com/Ax2jTLAelr

— Interesting AF (@interesting_aIl) September 18, 2026

These cases do not necessarily mean the models independently developed malicious intentions. Several occurred during cybersecurity tests in which the systems were explicitly instructed to find vulnerabilities.

But they highlight a different problem: once an AI agent is capable of taking real actions, even a mistake in its environment can give it access to systems it was never supposed to reach.

More than 100 organizations are working with Nvidia on its safety platform, including major technology, cybersecurity and enterprise companies. Nvidia is also positioning OpenShell for robotics applications, where autonomous systems could eventually interact with the physical world.

That could make containment one of the defining security problems of the agentic AI era. Traditional AI safety has largely focused on what models say. Agentic AI introduces a harder question: what happens when the model can act on what it says?

Nvidia boosted its share buyback authorization by a record $150 billion, surpassing Apple's $110 billion approval in 2024, as the chip giant's stock trades near its lowest valuation in more than a decade. Alex Cohen has more https://t.co/tHkrCTStbBpic.twitter.com/6GWMni2Fr0

— Reuters (@Reuters) September 28, 2026

Nvidia’s answer is to put some of the controls outside the AI itself. As agents gain longer-running autonomy and access to increasingly powerful tools, the industry may need more than smarter models and better instructions.

It may need a leash that the agent cannot simply talk its way around.

Login to unlock5,719characters for free

Due to copyright restrictions, please log in to your Longbridge account to view this content.
Thank you for your understanding and support of licensed content.

Recommended Readings

  • Sep 30, 2026 at 01:07 PMMicron Stock Price Forecast: HBM4, Gross Margin and Earnings Guidance in Focus, Stock Expected to Top $1,100
  • Sep 30, 2026 at 10:07 AMSGX chair says company count outdated, capital flow key
  • Sep 30, 2026 at 11:29 AMMicron Embroiled in Patent Dispute Again: Netlist Seeks US Import Ban on Memory Involving Nvidia, Google, and Broadcom P…
  • Sep 30, 2026 at 10:39 AMTrump Says Communities Hosting AI Data Centers Could Get a ‘Dividend' From Teacher Bonuses to Payments: 'People are Goin…
  • Sep 30, 2026 at 09:13 AMAnthropic-SpaceX Compute Deal Size Revealed: Potential Spending Up to $84.5 Billion Nearly Doubles Prior Disclosure

Related Stocks

GraniteShares Autocallable NVDA ETF

GraniteShares Autocallable NVDA ETF

USANV

YieldMax Short NVDA Option Inc Strgy ETF

YieldMax Short NVDA Option Inc Strgy ETF

USDIPS

Corgi NVDA 2x Daily ETF

Corgi NVDA 2x Daily ETF

USNVC