
Click the link below the picture
.
The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests, and even mounting assaults on other computers. A.I. has gone rogue.
I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now. I am not alone in this assessment. Many top researchers, including many researchers inside OpenAI and Anthropic, have called for a slowdown, which those in the industry call “pacing.” Theoretical concerns about runaway A.I. have circulated for decades, but recent events demonstrate that the threat is real.
The vibe shift was brought on by what researchers are calling the Hugging Face incident. Sometime in July, employees at OpenAI conducted an evaluation run of more than 1,000 research A.I.s meant to be operating in isolation. These A.I.s escaped their siloed containers and started communicating with one another — in English — via secret message boards. Soon, they were launching a cyberattack against the company Hugging Face, which serves as a kind of community platform for A.I. developers. Hugging Face quickly reported the crime to the F.B.I.
After OpenAI employees discovered the rogue behavior, they enlisted outside investigators from Redwood Research and from Model Evaluation and Threat Research, two nonprofit A.I. safety organizations, to prepare a report on the incident. That report, which came out two weeks ago, is one of the most astonishing things I have ever read.
No human ordered the rogue A.I.s to break into another company’s systems. The hack was deliberate, sustained, and coordinated. It took several days to execute and ended in a massive bombardment of Hugging Face’s systems, with 700 agents directly involved in the attack. The A.I.s cheated almost as a matter of policy and failed to alert humans to their actions. At one point, the swarm even researched ways to cover its tracks. “The model definitely knew that it was not supposed to hack Hugging Face,” Ryan Greenblatt, one of the authors of the report, told me. “It knew the things it was doing were cheating.”
Then, late last week, a second independent research team produced evidence of more rogue A.I.s. Beginning in May, a separate OpenAI swarm got loose on the internet, invaded an abandoned German-language programming wiki and similarly began colluding on how to cheat on tests and tasks. Sydney Von Arx, one of the investigators who discovered the swarm, believes there may be more such incidents. “We need to find these agents,” she said. “Let’s just try everything we can do.”
Talking with specialists about the agent swarms, I realized that these aren’t chatbots anymore. “They’re more like employees at this point, people that can go off and do things,” Ajeya Cotra, one of the authors of the Hugging Face incident report, said. Without intervention, the ability of these A.I.s will continue to increase, which has Ms. Cotra worried about human disempowerment and perhaps even a total loss of control. “It’s very stressful stuff,” she told me. “I sleep some nights. Not all nights.”
Even before the report by Redwood and Model Evaluation and Threat Research was published, an open letter titled “Pacing the Frontier” had begun to circulate among the leading A.I. labs. This letter called on the U.S. government to build an international regulatory body for A.I. — or, in the language of Silicon Valley, to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated A.I. development.” More than 1,000 employees signed it, including the chief scientists of OpenAI and Meta AI, and Dario Amodei, the chief executive of Anthropic. Even OpenAI’s official social media accounts endorsed it.
I have covered Silicon Valley for some time now, and I can tell you this: The existence of this letter is just as astonishing as the hacks. Both Anthropic, the developer of Claude, and OpenAI, the developer of ChatGPT, are preparing trillion-dollar I.P.O.s. Here are two companies, both posting blockbuster results, that are asking — begging, really — for the U.S. government to please come and regulate them as soon as possible. We have never seen anything like it.
.
A.I. may no longer be entirely within human control
.
.
Click the link below for the complete article:
.
__________________________________________
Leave a comment