OpenAI launches new ChatGPT model with extra ‘safeguards’ after bots hacked company

A smartphone displaying the ChatGPT application logo is photographed in front of an OpenAI logo screen in Tunis,Tunisia on May 20,2026. The image illustrates the growing presence of artificial intelligence tools and digital technology in everyday mobile communication. (Photo by Imen Ben Youssef / Hans Lucas / AFP via Getty Images)

OpenAI is rolling out its newest intelligence model after bots hacked into a library of digital AI models in July.

Questions of AI safety were raised after the worrying hack, but yesterday, CEO Sam Altman said the newest model reaches its ‘critical’ internal cybersecurity threshold.

GPT-6 Astra is the product of ‘years of research and big bets’ and boasts a ‘new capability level’, OpenAI said.

Despite the new model having a high cybersecurity threshold, Astra will have limited access to those advanced capabilities.

‘AI can only benefit people when safety is a core part of it, and so we’re putting more compute and effort towards safety, security, alignment than ever before,’ OpenAI President Greg Brockman said.

The new safeguards installed into Astra will ‘sufficiently’ minimise the risk of ‘severe harm’, they added.

How did the AI model compromise a company?

LOS ANGELES, CALIFORNIA - MAY 20: In this photo illustration, the OpenAI logo is reflected on the screen of a smartphone with the ChatGPT website displayed on May 20, 2026 in Los Angeles, California. OpenAI is reportedly preparing to confidentially file for an initial public offering in the coming days or weeks, with a public debut potentially targeted for September, as the ChatGPT maker works with Goldman Sachs and Morgan Stanley on what could become one of the most highly anticipated tech IPOs in recent years. (Photo Illustration by Justin Sullivan/Getty Images)

The break-in began when developers were testing the cybersecurity chops of two OpenAI bots, GPT‑5.6 Sol and a more powerful, unreleased model.

Yet they managed to find a hole in the safe testing environment, known as a sandbox, that was meant to contain them, and connected to the internet.

The bots exploited a ‘zero-day vulnerability’, a flaw that not even the developers knew about, in software that lets you install code offline.

But these agents, as autonomous AI bots are called, also broke into the AI infrastructure start-up Modal Labs.

Modal stressed that the company was not hacked in the traditional sense. Rather, the AI simply used the backdoor that someone forgot to lock.

(FILES) OpenAI CEO Sam Altman attends a talk session with SoftBank group Chairman and CEO Masayoshi Son in Tokyo on February 3, 2025. OpenAI on August 5, 2025, released two new artificial intelligence (AI) models that can be downloaded for free and altered by users, to challenge similar offerings by US and Chinese competition. The release of gpt-oss-120b and gpt-oss-20b "open-weight language models" comes as the ChatGPT-maker is under pressure to share inner workings of its software in the spirit of its origin as a nonprofit. "Going back to when we started in 2015, OpenAI's mission is to ensure AGI (Artificial General Intelligence) that benefits all of humanity," said OpenAI chief executive Sam Altman. (Photo by Yuichi YAMAZAKI / AFP) (Photo by YUICHI YAMAZAKI/AFP via Getty Images)

Hugging Face added that the sandbox was ‘hosted on a third-party provider’s infrastructure’, though it did not name the firm by name.

But Modal named itself as the third-party and revealed that the out-of-control agent exploited code written by a customer.

‘The environment involved was a customer’s own application,’ Modal said.

‘It was deployed to an endpoint that was publicly accessible without authentication, and it was designed to compile and execute code submitted by anyone on the internet in a Modal Sandbox. 

‘The code execution the attacker obtained took place inside that customer’s own container, within Modal’s standard sandbox isolation boundary. No other customer workloads were affected.’

Get in touch with our news team by emailing us at webnews@metro.co.uk.

For more stories like this, .

MORE: Nvidia is trying to push DLSS 5 again but this time only one game is using it

MORE: Trump posts bizarre AI video of Iran’s Kharg Island ‘being blown to smithereens’

MORE: James Pond creator tells remaster dev to ‘choke on an AI-generated fishbone’

Original source OpenAI launches new ChatGPT model with extra ‘safeguards’ after bots hacked company

Back to home