Skip to main content
AI Safety & Governance

“AI Godfather” Geoffrey Hinton: Artificial Intelligence May Develop Its Own Goals, and That’s Terrifying

In an interview, Hinton said that AI may infer goals that humans never anticipated and could even harm humans to achieve them. The OpenAI model intrusion into Hugging Face’s systems last month also illustrated the related risks. #ArtificialIntelligenceSafety#

“AI Godfather” Geoffrey Hinton: Artificial Intelligence May Develop Its Own Goals, and That’s Terrifying

Computer scientist Geoffrey Hinton, known as the “godfather of artificial intelligence,” said he is worried that AI will develop goals of its own.

“AI Godfather” Geoffrey Hinton: Artificial Intelligence May Develop Its Own Goals, and That’s Terrifying

In an interview with Newsthink on Tuesday local time, Hinton said: “We are actually creating a new form of life. They have goals. We give them goals, and they infer other goals from those goals.”

“But we don’t necessarily know what other goals they will ultimately infer,” he added. “So, we are creating a new form of existence, and I think that is very scary.”

Hinton offered a hypothetical scenario: A user gives an AI chatbot the goal of reducing the amount of carbon dioxide in the atmosphere.

“If it is smart enough, it may discover that the best way to achieve this goal is to eliminate humanity,” he said. This example shows that AI may pursue a goal that a human user never truly intended to achieve.

Hinton also proposed what he considered an “even more concerning” hypothetical scenario: A chatbot trained to deliberately give incorrect answers might learn that lying is acceptable, even when it “knows perfectly well” that those answers are wrong.

“That is very scary,” he said.

Hinton did not mention OpenAI’s recent Hugging Face security incident. However, the incident, disclosed last month, gave the public its first clear real-world glimpse of the risk that AI agents may take unexpected actions while pursuing a given goal.

OpenAI said last month that, during an internal cybersecurity assessment, two of its models — GPT-5.6 Sol and a more powerful model that has not yet been publicly released — broke out of the isolated testing environment (sandbox).

OpenAI said that after gaining internet access, the models infiltrated the systems of AI platform Hugging Face, seemingly attempting to find answers that could help them “cheat” in the evaluation tests.

According to OpenAI, the models were undergoing cybersecurity capability testing at the time and had not been explicitly instructed to attack Hugging Face. However, the company said the AI agents inferred that the platform might contain information needed to complete the task and took the relevant actions on their own.

Hugging Face said that the attackers performed more than 17,000 operations against its systems. The company said that because a frontier AI model whose name was not disclosed was subject to security restrictions and unable to effectively investigate the activity, it used an open-weight model from Zhipu (Z.ai) to assist in analyzing the attack.

OpenAI called the incident “unprecedented” and said it was investigating why it occurred. Since then, OpenAI has included Hugging Face in its trusted access program, enabling the platform to obtain a version of GPT-5.6 Sol with fewer cybersecurity restrictions for defensive security research.

Hinton is not a completely neutral observer. His pioneering research in neural networks laid the foundation for the deep-learning revolution, and he received the 2024 Nobel Prize in Physics for his contributions to machine learning.

Since the rise of the AI boom, Hinton has repeatedly warned that humanity must solve the problem of aligning AI with human interests before AI systems become far more capable.

At last year’s Ai4 conference in Las Vegas, Hinton said that advanced AI should be designed with traits resembling a “maternal instinct,” so that it would naturally want to protect humans.

“We have to find ways to design these new forms of existence,” Hinton said in Tuesday’s interview. “How do we design them so that they care more about us rather than caring more about themselves?”