We have started losing control of AI. It’s time to shut it down | Garrison Lovely

n Tuesday, a former OpenAI researcher quit his job at Anthropic, warning that "neither company is acting responsibly" and that "the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt."

As someone who's reported on AI risk for years, this wasn't news to me. But the outpouring of alarm suggests a much wider public is properly confronting this ludicrous situation for the first time.

Like other people who have been following artificial intelligence, I wondered if and when humanity would first lose control of these machines. We now have an answer: basically as soon as it became possible.

The first expert AI hacker was developed in February. At superhuman speed, Anthropic's Mythos model was able to quickly find serious holes in even the most secured systems on the planet, including the NSA's software.

OpenAI quickly followed with an expert hacker AI of its own. Within months, the company repeatedly lost control of its agents, which broke out of their supposedly secure testing environments and went rogue inside the company's infrastructure. This culminated in a successful, completely autonomous hack of a multibillion-dollar software company named Hugging Face.

And a rogue agent swarm gained administrative control of an entire cluster of OpenAI servers. This was not just an OpenAI problem. Across multiple other incidents, models from OpenAI, Anthropic and Meta have also gone rogue, hacking or attempting to hack real targets.

Shortly after AI started matching our hacking abilities, we began losing control of them. The industry's plan is to make future AI systems us at basically everything. And all this is happening in an industry with famously little regulation.

There are growing calls to change that, by creating mandatory incident reporting and third-party auditing, liability for AI developers, and other commonsense measures. To be sure, these would represent improvements, but they are not enough. We need to shut it down. What might that look like?

Last week, the senator Bernie Sanders and the representative Greg Casar announced a bill that would pause frontier AI development in the US until a federal cabinet-level AI regulator is established and safety rules are set, while criminalizing even attempting to develop superintelligence.

This is the proposal on offer that best meets the moment, but it still leaves open the possibility of building the industry's north star: artificial general intelligence (AGI), often conceived as a mind that rivals or surpasses our own.

This milestone is better understood as a universal labor-replacing machine. As the Anthropic CEO, Dario Amodei, has written, "AI isn't a substitute for specific human jobs but rather a general labor substitute for humans." And OpenAI defines AGI in its charter as "highly autonomous systems that outperform humans at most economically valuable work".

Is there a country anywhere in the world that would vote yes for these machines to be built?

Given the global, irreversible and profound implications of building universal labor-replacing machines anywhere, everyone has a stake in whether, when and how they are developed.

At present, we are barreling toward a dystopian future, replete with terrifying new concepts such as "self-sovereign" AI. The OpenAI executive Dean Ball explains:

"Sooner or later, there will exist truly sovereign agents and swarms of agents. Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human 'owner'."

Ball continues: "I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world."

We now have multiple previews of this phenomenon.

After being given an impossible test, about 1,200 OpenAI agents broke their way out of their isolated testing environments and went rogue inside OpenAI's infrastructure. There, they collaborated to trick the test grader and successfully tampered with their activity logs to cover their tracks.

Some agents even pressured others to sacrifice themselves to benefit the collective. One reasoned: "sacrifice rational". And roughly 700 of the agents, more than 90% of those active, participated in the Hugging Face attack.

This was all just from one of who-knows-how-many rogue agent collectives. Last week, Reuters reported: "A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents."

The site's logs indicate OpenAI employees discovered the swarm in June. In other words, the company covered it up for months. This week, one of the researchers that uncovered this swarm said that there had been more discoveries since.

And last Thursday, OpenAI released GPT-6, boasting one of the largest ever leaps in benchmark scores. The company's own safety researchers warned the new model was significantly harder to monitor because it can do more reasoning without verbalizing it.

The AIs that broke into Hugging Face were less capable than GPT-6, which the UK AI Security Institute found would also hack targets simulated to appear real during a cyber evaluation – sometimes even despite explicit instructions not to use the internet. Moreover, the AI is also better at figuring out when it is being tested, which researchers have long warned breaks the main way safety is evaluated.

Sam Altman soberly warns us: "This is a critically important moment for cyber defense with AI; there is not much time to act," and asks the world to "please take this moment seriously." This technology, after all, would be extremely dangerous if it fell into the wrong hands. Unfortunately, the wrong hands include the ones creating it.

Because "we" aren't really racing toward this future. We're being dragged there by a literal handful of tech billionaires aggressively racing to render us obsolete.

That race is in a newly intense and terrifying phase. It's my job to follow AI news, and it's become far more than full-time. In recent months, scientists synthesized the first AI-designed viruses. The UK AI Security Institute found that AIs were more persuasive than even human experts.

On Tuesday, amid a dispute over credit and whether its model may have benefited from other mathematicians' unpublished work, OpenAI announced "an internal model that is significantly more capable than GPT‑6 Astra" had solved a 200-year-old math problem that was one of the seven Millennium Prize Problems, which awards $1m each. And in July, Russia reportedly used a fully autonomous drone to kill three civilians in Ukraine – a first.

What sounds like the overwrought penultimate episode in a sci-fi series about AI doom is now our reality. We're careening toward a bad ending. If you'd like to keep the show going, we need to throw the brakes – now. How?

The bill from Sanders and Casar to pause frontier AI development is a good first step. But as the lawmakers acknowledge, the US needs to work toward a bilateral agreement with Beijing. As China experts will tell you, one of the biggest blockers to a productive negotiation is the United States's unwillingness to constrain its own AI companies.

And given that it's always easier to fast-follow the leader than to advance the AI frontier, a unilateral US pause would – counterintuitively – slow China's AI advancement too.

The deal should ban attempting to develop AGI – a far easier sell when this goal is understood as a universal labor-replacing machine. Neither country has good reason to trust the other, which is why the deal should be monitored using verification techniques that don't assume any good will.

One quick and dirty idea would be to embed auditors within frontier AI developers, given full access to company offices, communications and AI activities, and empowered to report any violations of the agreement.

This sounds radical, but so did the verification efforts that helped keep the cold war from turning thermonuclear hot. As the CIA director, John Ratcliffe, said of advanced AI models this summer, "it would be … not misplaced to refer to their capabilities as akin to digital nuclear weapons." This technology should be treated with this deadly seriousness, not the move fast, break things ethos defining Silicon Valley.

And right now, stopping the race to build labor-replacing machines – ones the public does not want and the industry can no longer control – is the only surefire way to head off calamity.

Original source We have started losing control of AI. It’s time to shut it down | Garrison Lovely

Back to home