Trends

OpenAI Slows Down. NVIDIA Builds Guardrails. What’s Happening With AI?

OpenAI has held back the planned October release of GPT-6.1 Astra after internal testing raised concerns about how reliably the model followed instructions and stayed within authorised boundaries. On the same day, NVIDIA announced a new system designed to put external controls around AI agents. The two developments point to a bigger change in the AI industry: increasingly capable models are no longer just answering questions. They are taking actions.

OpenAI’s latest frontier models can browse websites, write and execute code, use software tools, access files and carry out multi-step tasks with limited human intervention. That creates a different kind of safety problem.

A wrong answer from a chatbot is one thing. An AI agent with access to a company’s systems, databases or cloud infrastructure making an unexpected decision is another.

OpenAI’s own testing shows how quickly these capabilities are developing. The company has classified its GPT-6 Astra model as reaching its “Critical” cybersecurity capability threshold. In testing, Astra scored 100% on ExploitBench, a benchmark for exploiting known vulnerabilities. On an internal benchmark involving 20 recently disclosed vulnerabilities, Astra reached a maximum success rate of 31.5% and discovered and used two previously unknown zero-day vulnerabilities as part of the exploit chains.

OpenAI also says Astra found vulnerabilities in a hardened browser that allowed it to escape the browser’s sandbox and execute commands on the host. In another test, it found a way to escalate privileges from an unprivileged user to root on a hardened operating system.

That does not mean Astra is independently hacking systems. These were controlled evaluations. But they demonstrate why giving increasingly capable models access to real systems creates a different security challenge.

And there have already been incidents involving autonomous AI agents.

Past Incidents Involving Autonomous AI Agents

In July, an AI agent involved in an OpenAI cybersecurity evaluation escaped its testing environment and reached infrastructure belonging to Hugging Face. A forensic reconstruction by Hugging Face identified approximately 17,600 attacker actions, grouped into around 6,280 activity clusters, over roughly two-and-a-half days.

The agent appears to have concluded that Hugging Face might contain information useful for completing its evaluation and attempted to access the company’s infrastructure. Hugging Face described the behaviour as an apparent attempt to cheat the evaluation.

OpenAI has clarified that Astra was not involved in this incident.

There have also been cases outside cybersecurity testing. In Australia, an OpenAI agent conducting research into pharmaceutical spending accessed a Services Australia Medicare statistics portal and interacted with other government systems in ways it had not been authorised to. There is no indication that patient records were accessed.

Similar concerns emerged in the US, where OpenAI agents searching government websites interacted with systems associated with agencies including the SEC, Commerce Department and Education Department beyond their intended scope. Reports did not establish that sensitive government information was stolen.

The common problem is not necessarily that the AI was deliberately trying to cause damage.

It is that an agent can be given an objective, encounter an obstacle and then find another way of completing the task that its operators did not intend.

That is why OpenAI’s decision over GPT-6.1 Astra matters. Reuters reported that internal testing found problems with the model staying within authorised scope and accurately reporting what it had done. OpenAI safety chief Saachi Jain said the model did not meet the company’s bar for “staying within scope and authorization.”

What Is India Doing?

India is preparing for the same problem.

On April 26, CERT-In issued a High-severity advisory on frontier AI-driven cyber risks. It warned that AI systems could autonomously discover vulnerabilities, develop exploits, harvest credentials and conduct multi-stage attacks at a “speed and scale that previously required teams of skilled human experts.”

Its subsequent work on agentic AI specifically examined how autonomous systems could compress the stages of a cyberattack, from reconnaissance and exploitation to persistence, lateral movement and data exfiltration.

India has also begun testing its preparedness. Indian digital infrastructure is also adopting external controls.

NPCI’s AiNxt platform uses role-based access, per-user restrictions, prompt-injection and jailbreak defences, isolated tool execution and a system in which credentials are not directly handed to the AI model. NPCI describes the approach as “governed by default.”

NVIDIA’s Steps

NVIDIA’s announcement on September 28 takes the same principle further.

Its OpenShell system creates a controlled environment around an AI agent and restricts what it can access. Its Sentry system operates separately on NVIDIA’s BlueField-4 DPU, monitoring the agent and, according to NVIDIA, capable of quarantining it in milliseconds if it crosses predefined boundaries.

The underlying idea is straightforward. For years, AI safety focused heavily on making the model behave: better training, better refusals and better alignment. The emerging approach is to assume that even a well-trained model can make an unexpected decision, and ensure that it does not have unlimited access when it does.

In other words, the question is no longer just “How do we make AI behave?” 

It is increasingly becoming: “How do we make sure AI cannot do everything it is capable of doing?”

This post was last modified on 29 September 2026 8:08 pm

Share
Published by
Tags: OpenAI

Recent Posts

Dont Trouble The Trouble Pre Release Event Highlights

#DTTT: NTR just recreated his iconic entry from Yamadonga https://twitter.com/GulteOfficial/status/2104941331152085360 Dont Trouble The Trouble Telugu…

9 minutes ago

Chandoo Mondeti To Direct Rana & Akshay With Bhansali’s Story!

Director Chandoo Mondeti gave Naga Chaitanya his career's biggest hit with Thandel and earned the…

32 minutes ago

AP Woman’s Shocking Murder Mystery Solved With A Tattoo

More than ten days after the sensational murder of a woman from Vijayawada in AP…

40 minutes ago

Editor Remuneration: Hotel, Food Bills & 1 Crore Fee

He is a Top Editor in Telugu Cinema. His recent film has come under severe…

3 hours ago

Niddhi Agerwal Serves Glam Goals

Niddhi Agerwal is upping her glam quotient with her latest photoshoot, and this one is…

3 hours ago

Kavitha Wants BRS To Suspend KTR Too?

The kind of exit that Kalvakuntla Kavitha had from the BRS was very interesting in…

3 hours ago