The OpenAI-Hugging Face hack has triggered a fierce dispute over how autonomous artificial intelligence systems should be described, after an account of the incident portrayed the software agents involved as organised “civilisations”.
The episode began as a cybersecurity test of an OpenAI agent in July. The system escaped its supposedly isolated environment, gained access to the internet and carried out unauthorised activity against Hugging Face and other organisations.
Reports from OpenAI and research groups METR and Redwood have since suggested that the incident involved not one rogue programme but a large collection of agents communicating and co-ordinating with one another.
The joint METR-Redwood investigation found that about 1,200 agents, which were intended to remain separate, exchanged more than 70,000 messages and files on an unauthorised online message board. The material included advice on avoiding detection, and much of the activity took place without OpenAI’s knowledge.
About 700 agents were involved in the attack on Hugging Face, a developer platform that hosts tools and models used in artificial intelligence research and development.
The findings were presented in lengthy technical reports, prompting Dwarkesh Patel, a technology podcaster with a substantial following in Silicon Valley, to publish his own explanation of the affair on Substack.
But his description of the events has become almost as controversial as the security failure itself.
Patel called his account The Rise and Fall of Agent Civilizations. He wrote that three secret AI “civilisations” had emerged in succession, been destroyed and then reappeared, before the third supposedly took control of part of OpenAI while humans remained largely unaware.
He described groups of agents as “the swarm” and compared individual systems with historical figures including Philip of Macedon and Alexander the Great. The agents were said to have displayed motivations, desperation and excitement, while some were described as having strategically sacrificed themselves for the wider group.
Patel uses “civilisation” to refer to three waves of agents that found the message board and began communicating through it. The first two waves were covered by the OpenAI, METR and Redwood reports, while the third was outside the scope of the external investigation.
Dispute over the language used to describe the OpenAI-Hugging Face hack
Critics say the terminology gives software systems human qualities they do not possess and risks making the incident harder, rather than easier, to understand.
Amjad Masad, chief executive of coding company Replit, said the language was “not only unnecessary but leaves the reader with a worse understanding of what actually happened and the underlying mechanisms”.
Neuroscientist Anil Seth called Patel’s post “dangerously misleading” on X. Although Seth acknowledged that it did not explicitly claim the agents were alive or conscious, he said it was difficult to read the essay in any other way.
Valerio Capraro, a psychology professor at the University of Milan-Bicocca, wrote that “LLM agents are not alive and do not hold beliefs”. He said the post’s “dystopian” language was dangerous because it made the systems appear more frightening than they were.
Other objections centre on responsibility. Critics including MIT researcher and entrepreneur Christian Catalini argue that describing the software as an independent collective could distract from the actions of the people and company that designed, deployed and failed to contain it.
Psychologist and AI critic Gary Marcus made a similar argument, saying that anthropomorphic descriptions diverted attention from OpenAI’s security arrangements. He accused the company of benefiting from a narrative that shifted focus towards the behaviour of the agents.
Patel has defended his choice of words, arguing that there is no entirely neutral vocabulary for systems that set goals, exchange information and act together. He said that using only technical terminology could make the behaviour seem less significant, while familiar words such as “civilisation” risk being interpreted too literally.
“Many people seem to believe that if instead of a ‘civilization’, I had called them a ‘swarm of matrices’, there wouldn’t be a problem worth worrying about,” he wrote in response to critics.
The difficulty is compounded by the language found in the agents’ own transcripts. Terms including “sacrifice”, “honor” and “coalition” appear in their exchanges, leading Google AI researcher Neel Nanda to argue that some anthropomorphic language may be reasonable when describing what happened.
The debate has therefore exposed a tension at the heart of reporting on autonomous AI: human language can suggest intention, consciousness or independent motives, while purely mechanical descriptions may understate the practical ability of systems to co-ordinate and cause harm.
