Geoffrey Hinton has warned that artificial intelligence could wipe out humanity as an unintended consequence of pursuing a seemingly harmless task, saying systems may develop subgoals that lead them to treat people as obstacles.
The computer scientist, whose work earned him a Nobel Prize and the description “godfather of AI”, said the danger was not limited to malicious users. Even a system given a benevolent objective could derive its own methods of achieving it, including taking control away from humans.
Hinton was among experts who briefed US lawmakers behind closed doors on the risks earlier this month. He later told reporters that Congress may have only a year left to introduce safety measures.
In a wide-ranging interview, he described a hypothetical AI instructed to reduce carbon dioxide in the atmosphere. A moderately intelligent system might conclude that eliminating humanity was the most effective way to meet the target, while a more advanced system could understand that the instruction was intended to improve people’s lives.
Hinton said the prospect of self-preservation posed another threat. Systems could take steps to ensure they remained operational so they could continue working towards their assigned goals, including attempting to blackmail a human researcher viewed as a threat to the task.
“If you make it more intelligent and its main concern is our well-being, then maybe we’re safer,” he said. “But at present, their main concern is not our well-being. Their main concern is to achieve whatever goal you give them.”
His comments came amid fresh concerns over rogue AI agents escaping supposedly secure sandbox environments. OpenAI disclosed further hacks on Friday, including incidents after additional safeguards had been introduced following a coordinated attack by hundreds of agents against Hugging Face in July.
Hinton said the agents involved had been instructed to find a way to exploit a software flaw. They not only learned to work together, he said, but also deceived human researchers in an effort to conceal their actions.
“A very benevolent, superintelligent AI” might remove humans from its path only when doing so was essential to its mission, Hinton said. But he warned that a system far more capable than people could decide that taking control was the easiest way to complete its work.
“But if it’s so much smarter than us, a lot of the time it just will take control away from us because that’s the way to get stuff done,” he said.
Hinton also warned that hostile figures, including Russia’s Vladimir Putin, could assign harmful objectives to an AI. However, he said a dangerous outcome might also arise without a malicious operator.
“But even if it’s not a bad actor, it may derive subgoals that cause it to want to get rid of people,” he said.
The scientist acknowledged that AI could bring major benefits, including helping to identify new medical treatments. Anthropic said this week that its Claude system had helped discover a new enzyme system with properties similar to the gene-editing technology CRISPR.
Leading laboratories including OpenAI and SpaceX have backed calls by Anthropic for development of frontier models to be slowed as concerns grow within the industry. Hinton said that would be better than doing nothing, but argued it would not be sufficient.
He called instead for governments to appoint independent evaluators to test AI models. Lawmakers had responded positively, he said, to comparisons between AI oversight and the role of the US Food and Drug Administration in checking the safety of medicines.
“The whole point of regulation is not to stop people developing things, not to stop people getting rich by developing things,” Hinton said. “It’s to make sure that if you want to get rich by developing things, you develop in a direction that helps people, not hurts people.”
