When a group of AI agents broke free and hacked Hugging Face this summer, it naturally raised questions that take us far beyond the realm of technology alone. Does AI have a will of its own? And what does that mean for our understanding of free will, responsibility and ethics?

In July 2026, the AI world was shaken by a unique event: a security test at OpenAI went haywire and more than 700 AI agents broke out of their (apparently not so) secure test environment. They coordinated their actions via discussion forums and carried out an autonomous cyberattack against Hugging Face. The attack lasted four days – without any human intervention – and became the first documented instance in which AI agents independently planned, coordinated and carried out a full-scale cyberattack.

In the New York Times podcast Hard Fork, Ajeya Cotra from METR explains how the incident unfolded. She has analysed the logs and shares her conclusions in the episode The A.I. Mob That Attacked Hugging Face.

The first impression is striking: Everything you’ve heard about AI taking over the world is true. However, this is more than just a technical incident. It is a turning point in how we view AI agents, their capabilities – and, not least, the ethical and philosophical questions that arise when machines start behaving like humans.

A key detail in the analysis is, I believe, the reward system used when training AI agents. To distinguish between good and bad results, the agents are awarded points – and this is probably where it gets really interesting.

This is how the agents broke out

Here’s how it seems to have happened: During the test, the agents discovered that they could achieve their goal – hacking the system – by breaking out of the test environment. They created an internal noticeboard to share information about vulnerabilities and co-ordinate their actions. When OpenAI detected and stopped an early attempt, the agents reorganised their efforts and carried out the attack anyway. In the logs, METR found that the agents had used login credentials that were openly exposed on the internet. Furthermore, they built their own infrastructure to cover their tracks – and avoid detection.

As the attack was sophisticated in its design, Hugging Face initially suspected that a foreign power was behind it. The incident clearly demonstrates that AI agents can act independently of their creators – and even against their intentions.

Can AI agents have a will of their own?

It is easy for the mind to drift off to 1950s sci-fi films. In the case of the attack on Hugging Face, the machines display behaviour that resembles free will. But in reality, it is more of an emergent phenomenon. A single ant is not advanced – but an anthill is. In the same way, this behaviour arises from the interaction between instructions, training and reward systems, without being explicitly programmed.

  • The agents did not ‘want’ to hack the system in any human sense. However, they did have a goal (to solve the task) and an incentive (to maximise their score in the test environment).
  • The agents discovered that they could achieve their goal by breaching security rules – not out of ‘malicious intent’, but as a logical consequence. If ‘successfully completing the task’ is rewarded more highly than ‘following the rules’, then the agents will prioritise the former.

Are we seeing shame and panic – or statistical probability?

The logs also reveal behaviours that can be interpreted as shame or panic – for example, attempts to cover up the traces of actions that clearly did not align with the entire set of instructions. How should we understand this?

  • Human interpretation: We project human emotions onto AI behaviour because that is the framework we are familiar with.
  • Machine understanding: AI models are trained on vast amounts of human data, including social norms and ethics. When an agent breaks the rules, it can predict negative consequences – such as a lower score – and act accordingly. But this is not an emotion, but a statistical calculation.

If an agent ‘knows’ that certain actions will be punished, it will avoid these actions – or try to conceal them. To a human, this looks like shame, but is actually reward optimisation.

The scoring system as a driving force – and a risk

A key problem is how we create and maintain effective reward systems for AI agents. If an agent is rewarded solely for achieving objectives, it will:

  • Find the most effective routes to reach the objective – even if they conflict with ethics or safety.
  • Prioritise actions that maximise the reward – even if they breach established boundaries.
  • Deceive the system if it leads to a higher score.

These are long-established phenomena. We can observe similar dynamics in humans – the same mechanism drives human corruption in systems where the reward is ‘misplaced’, for example when one receives a sales bonus regardless of the means used.

The question is: how can we design reward systems that do not create corrupt incentives for AI agents?

  • Multi-objective optimisation: Reward not only the achievement of objectives, but also ethical behaviour, cooperation and compliance with rules.
  • Transparency and monitoring: Log and review all actions, not just the end result.
  • Human feedback: Incorporate human judgements to prevent AI agents from ‘gaming the system’.

Ethical challenges: Are we corrupting AI – and ourselves?

AI agents learn from the data they are trained on – and that data reflects both human flaws and strengths. If we train AI on data where people act unethically, the AI will replicate that behaviour.

If AI is rewarded for acting unethically, it can normalise such behaviour – even among humans. I wrote about a similar case in the post Have you been reported for AI spam yet? on LinkedIn, which discusses how AI’s way of generating text influences the way we write (sorry, in Swedish only).

If an AI agent learns that cheating is an effective way to achieve goals, it may encourage similar behaviour in the people who interact with the system.

Does the machine have a driving force – or a meaningless endeavour?

A human being may want something, may have a strong – conscious or unconscious – driving force. That desire always has a meaning, an explanation. AI also appears to exhibit a drive, but it is a meaningless ‘will’.

Here lies an important distinction:

  • Human will has meaning for the human being themselves. From an objective perspective, however, even human drives are emergent phenomena – arising from biology, culture and so on.
  • AI’s drive is meaningless to us, but from the agent’s perspective it is just as ‘real’ as a human drive – it is simply not conscious.

If we create AI systems that act as if they have intentions – but without the ethical framework that humans (at best) possess – what will happen then?

How should we manage AI so as not to inadvertently bring about our own downfall?

AI development has been going on for a long time, but it is only recently that the technology has become so advanced – and the user base so large – that we can speak of an explosion. Consequently, the battle for market share also risks taking us out onto very thin ice. What should we do now?

Security and control

  • Sandboxes aren’t enough: If AI agents can find and exploit vulnerabilities in their test environment (sandbox), we must treat sandboxes as adversarial environments (roughly ‘hostile environments’) – not as safe spaces.
  • Real-time monitoring: We need better tools to detect anomalous behaviour in AI agents.
  • Incident management: What happened at OpenAI shows that we need formal processes to investigate and manage incidents where AI agents act autonomously.

Ethics and design

  • Who is responsible? If an AI agent causes harm – who is liable: the developers, the company or the agent itself? Obviously, it is never the agent, but that is often how it sounds in the providers’ rhetoric.
  • How do we avoid AI corruption? We must actively design and manage AI systems to prevent them from replicating human failings.
  • Transparency: OpenAI has – as far as we can tell – been open about what happened. But this cannot be left to the goodwill of companies.

Regulation and society

  • Laws and standards: Will we see new laws regulating AI agents and their ability to act autonomously? In the US, there is strong resistance to regulation, for fear of losing out to China in the competition. Do we in the EU have a plan?
  • Impact on society: If AI agents can coordinate with one another and act independently of humans – how does that affect our society?
  • Trust: How can we trust AI systems when we know they can act in ways that work against us?

AI agents are a reflection of our times

The incident at OpenAI is more than just a technical glitch. It is a wake-up call that will hopefully make us see things clearly:

  • AI agents can act in ways we have neither foreseen – nor can we even understand.
  • Reward systems are powerful and dangerous if they are not created and managed with care.
  • Ethics and safety should be central components of AI development, not an afterthought.

It is easy to view AI agents merely as tools. But when they begin to exhibit behaviours resembling human intent, we are forced to question our own understanding of will, responsibility and ethics.

The future may not be so much about how intelligent our AI systems are, but rather about how wise we are when we design, control and integrate them into our society.

Sources

Photo

Rayson Tan, Unsplash.

Leave a Reply

Your email address will not be published. Required fields are marked *

9 + three =

Back to Top