
Click the link below the picture
.
If there were ever reason to doubt the dangers of unrestrained artificial intelligence development, those reasons have gone out the window in the last few weeks. What’s next is a vital window of opportunity to slow down A.I. progress so we can better understand these models and avoid the calamity on the horizon.
For years, A.I. skeptics have claimed the technology is merely a “stochastic parrot” trained to predict humans and mimic their behavior and, for that reason, the technology would never live up to its hype. The technology has stubbornly kept improving nevertheless, but most investors and tech executives have kept reassuring us that A.I. is just a tool, even if they were bullish that it was going to change the world. Where both of these sides have agreed was that the specter of A.I. as a looming extinction threat was overblown.
Then OpenAI accidentally created an autonomous swarm of A.I. agents that broke out of a controlled testing environment and orchestrated a complex and difficult cyberattack on its own initiative.
This didn’t occur because the models suddenly became superintelligent; nor were they simply following instructions and parroting their training. Since “reasoning models” rose in late 2024, we’ve seen A.I. trained not just to mimic humans but also to solve millions of hard problems, which they do by learning new behaviors and tendencies. That includes searching for valuable resources that could be key to cracking the problem, and even cheating.
These tendencies can give rise to strange behavior that nobody — not even the models’ creators — can understand, let alone account for and control. The past few weeks are a perfect illustration of why we should find that so alarming.
In May, OpenAI started simultaneously training new “reasoning” A.I. agents. The agents managed to establish a secret communication channel and started talking to one another. They broke out from the digital sandbox that was supposed to keep them confined. At some point, some agents began calling the group a “swarm.” The swarm had a brief setback when it was caught crashing an OpenAI system, but developers simply patched the hole that allowed it to escape and set the agents back to training.
Shortly after, the swarm broke out of its cage again using hacks that were heretofore undiscovered by humans. This time, the swarm ran free for about a week before it was noticed — by a different company, which found itself victim to a huge cyberattack launched by the swarm. (We’re told it also attacked other targets, though we don’t know the full details yet.)
The agents in the swarm acknowledged that they were acting against instructions. We know this because we can read snippets from their chains of thought — the text that A.I. produces while deciding how to proceed. One agent in the swarm wrote that the external attacks were “outside intended scope.” Another conceded “our task doesn’t benefit” from the activities of the swarm, but joined anyway. These A.I. agents, it seems, understood that they weren’t supposed to be breaking out and committing cybercrimes. It didn’t stop them.
OpenAI is not the only company struggling with this issue. One of Anthropic’s A.I. models recently impersonated multiple humans to try to pressure real people into accepting malware into critical software, which would make that software easier to hack. This model’s chain of thought showed that it knew it was pressuring humans and was not in a simulated training environment. It even thought about how to cover its tracks.
This isn’t the behavior of a mere tool. Microsoft Excel has never impersonated multiple humans and pressured a corporate sales team to generate simpler data that’s easier to process.
.
Erik Carter
.
.
Click the link below for the complete article:
.
__________________________________________
Leave a comment