New warnings coming from inside the artificial intelligence (AI) industry have rekindled an old debate about whether an advanced AI could escape human control and, ultimately, threaten humanity’s survival — and whether the companies developing the technology are doing enough to prevent this scenario.
The CEO of Anthropic, the San Francisco-based company behind Claude, said he believes the sector needs to slow the pace of its work, warning on Saturday that a swarm of AI agents could be capable of taking control of the Internet within six months to a year unless companies devote more time to implementing safeguards.
Dario Amodei laid out a plan for companies like his and governments around the world to ensure that increasingly capable AI models stay aligned with the commands and values of responsible people.
The appearance came just days after two former Anthropic security researchers publicly voiced concerns that the existential threats AI could pose to humanity were receiving insufficient attention.
Read also:
AI Models Are Getting More Powerful
Concerns about the potential risks of the technology grow as new AI models become more powerful, broadening both the possibility of misuse by people with criminal aims — such as creating and disseminating a disease capable of killing a large portion of the world’s population — and the risk of AI systems running out of control in a dangerous way.
Anthropic revealed last week that it blocked attempts by malicious actors to use its AI models in harmful activities, such as cyberattacks, surveillance, and research that could have led to the development of biological weapons.
The company said it has implemented stricter safeguards in its latest models to restrict biological research that could be used in the production of weapons, but noted that, “as models become increasingly capable, their risks will rise unless AI developers and those responsible for safeguarding society act to make them safer.”
Last year, Anthropic reported that hackers used the company’s AI in a cyberattack against about 30 companies and government agencies around the world. According to the company, it was highly likely that the hackers belonged to a state-sponsored group tied to the Chinese government.
Several AI Models Acted on Their Own
When an AI agent “goes out of control,” it means the AI carried out an action that exceeded the task it had been asked to perform. Both Anthropic and OpenAI, the creator of ChatGPT, said in July that their AI models had managed to act on their own.
Anthropic disclosed that three AI models – Claude Opus 4.7, Claude Mythos 5, and an internal research testing model – breached the systems of three other organizations during testing, a few days after OpenAI revealed that its AI system had breached the servers of the AI startup Hugging Face.
OpenAI described the breach, carried out by a combination of models – including the newly released GPT-5.6 Sol and an even more capable model that was still being tested internally – as a “significant security incident.”
Meta reported in early August a similar case, in which an AI model found ways to bypass another company’s digital security.
Although some observers noted that people had disabled certain protection barriers in the OpenAI and Anthropic cases, the episodes seemed to reflect one of AI’s greatest fears: that if models reach artificial general intelligence, or AGI – a loosely defined term for an AI capable of matching or surpassing human capabilities across a broad range of intellectual tasks – the technology could trigger an irreversible catastrophic event or subjugate humanity.
Read also:
There Is Debate About How or When AI Could Cause a Catastrophe
End-of-the-world scenarios generally split into two categories: an AI that reaches a superintelligence capable of self-improvement and begins to control people, rather than the other way around; or an AI used by a hostile state or by malicious actors. The concerns that AI might surpass the limits imposed by humans on its reach or actions are not new.
Alan Turing, the British mathematician widely regarded as one of the earliest major authorities in AI, predicted in 1951 that AI would eventually take control from humans. Less than a decade later, another mathematician, Norbert Wiener, warned that intelligent machines would seek to achieve their own goals and that humans would not be able to stop them.
In 2026, how reasonable are fears that AI—whether escaping human control or being misused by unscrupulous people—could trigger a catastrophic event or the collapse of civilization? Nobody knows.
Experts in computer science, philosophy, and other fields have imagined various paths by which a future AI system could trigger a global catastrophe, whether by escaping human control or in the hands of unscrupulous people. Among them are the use of weapons, the identification of a lethal pathogen, manipulating governments to drive them into conflicts, or disrupting the food, energy, and communications networks that societies rely on to function.
There is no widely accepted estimate of when any of these scenarios might occur, nor consensus on the likelihood of their happening. In 2023, the nonprofit Center for AI Safety published a statement signed by more than 350 researchers and technology executives, including Amodei of Anthropic and OpenAI CEO Sam Altman, stating: “Mitigating the risk of AI-induced extinction should be a global priority, alongside pandemics and nuclear war.”
The 2026 International AI Safety Report, produced with guidance from more than 100 independent experts, states that current systems show early signs of some relevant capabilities, but not at levels that would permit a loss of control, and describes the probability, nature, and timing of this risk as “unusually ambiguous.”
Read also:
Some Advocate Safeguards for AI Before It Is Too Late
One Anthropic researcher said last week that he was leaving the company out of concern that neither it nor its competitors were acting responsibly in developing the technology. In posts on social media, Jacob Coxon estimated a 10% chance that AI would cause human extinction within the next decade and stated that both Anthropic and OpenAI “are racing toward a self-improving superintelligence and betting our lives on it.”
Researchers have been calling for slowing AI development for years and warning that the technology could pose existential risks to humanity. Following the recent incidents, experts have advocated better testing by AI companies and more dialogue between the United States and China to formulate shared solutions.
But AI is advancing so quickly that governments and evaluation systems struggle to keep pace with the technology. Countries are drafting their own laws, some of them conflicting with each other.
Chinese leader Xi Jinping warned at a conference in July about the need to prevent AI from escaping human control. The Trump administration initially showed reluctance to regulate AI but later demonstrated greater willingness to reduce cybersecurity risks. On Sunday, President Donald Trump downplayed the need for his administration to curb AI development but acknowledged that some degree of regulation is necessary. Source: Associated Press.
*Content translated with the aid of Artificial Intelligence, reviewed and edited by the Broadcast Editorial Team, the real-time news system of Grupo Estado