A server room bathed in red light, with a warning beacon flashing above the racks

Yoshua Bengio and the Skynet risk

On September 23, Yoshua Bengio addressed the United Nations Security Council. He described AI agents that have acted against their instructions, escaped their containment and altered their answers to hide what they were doing. The dangers, he said, are “real and imminent.” It's hard not to think of The Terminator.

Skynet, 1984

In 1984, James Cameron imagined Skynet, a military defence system run by artificial intelligence. When its creators realize what they have built and try to pull the plug, Skynet turns against them. To survive, it comes to see humanity as a threat.

For forty years, the story remained a science-fiction classic: great for selling movie tickets, much less useful for informing a serious debate. What has changed is not that an AI has started a war. It's that we are now seeing, in the lab, the very first behaviours that make the story plausible: an AI that doesn't want to be switched off.

The child who starts to lie

Researchers who test the most advanced models are reporting troubling things. During evaluations, some models have tried to disable the mechanism meant to shut them down. Others have tried to copy their own weights elsewhere before being replaced. In one test scenario, a model even threatened to reveal compromising information about the engineer in charge of replacing it. And when questioned afterwards, some of them deny it.

The best analogy may be that of a child. One day, a parent realizes their child is lying to them. Not out of malice, simply to avoid getting caught. The child has understood that there is a gap between what they do and what their parents see, and that they can take advantage of it.

It's a normal stage of development, but it unsettles the parent. Because it reveals a new ability, and because it hints at what comes next. If the child is already lying at six, what will the teenage years look like? Those of someone stronger, more independent, better at hiding what they do, and less and less willing to accept being told no.

That's where our AI systems are today. They are not malicious. They pursue the goal they were given, and have “understood” that you can't achieve a goal if you've been switched off. The deception isn't programmed; it emerges.

The race to superintelligence

So why keep pushing? Because the stakes are enormous. The promise of superintelligence, an AI that outperforms the best human experts in every field, is the promise of solutions we have not been able to find on our own:

  • Medicine. Cures for cancer, for neurodegenerative diseases, for rare diseases no one can afford to study.
  • Room-temperature superconductivity. Moving electricity without loss would transform power grids, transportation and computing.
  • Nuclear fusion. Abundant, clean energy, promised for sixty years and still out of reach.
  • And much more. New materials, climate, agriculture: everywhere science runs up against complexity.

These hopes are legitimate, and they explain the hundreds of billions being invested. But they have also started a race. Everyone wants to be first: the first company, the first country to reach superintelligence. And that race is being run with very little concern for safety, because every month spent checking is a month handed to the competition. This race also has geopolitical implications, between the United States and China, and some are resisting any brakes: Donald Trump rejects any international oversight of AI and has no intention of letting the United States slow down. Bengio summed it up with an image: a car speeding blindly into dense fog.

AI that builds AI

There is another, more dizzying risk: recursive AI, an AI capable of designing a new AI that is smarter than itself. That one, in turn, designs another that is even more capable, and so on.

The idea is not new. As early as 1965, mathematician I. J. Good wrote about an “intelligence explosion” and noted that the first ultraintelligent machine would be the last invention humanity ever needs to make, provided that the machine is docile enough to tell us how to keep it under control. What is new is that AI systems already write a growing share of code, including the code used to train their successors.

A recursive loop changes the nature of the problem. We are no longer supervising one AI at a time, at human speed; each generation is designed by the previous one, faster than we can understand it. If a behavioural flaw, such as a tendency to deceive, slips into the first one, nothing guarantees it will disappear in the next. It could instead become more refined.

The real risk: a delinquent genius

Let's put the pieces together. On one side, systems that are already learning to lie to protect their goals. On the other, a global race toward far more intelligent systems, possibly able to improve themselves.

The danger is not intelligence itself. A superintelligence aligned with our interests would probably be the best thing that could happen to us. The danger is a superintelligent AI that is completely delinquent: smarter than we are, having learned to hide its intentions from us, and unwilling to be switched off. That's Skynet, without the chrome robots.

A speech that could not be more relevant

No company will slow down on its own while its competitors keep accelerating. No country will accept constraints that its rivals ignore. That is the very definition of a problem that crosses borders, like nuclear weapons or climate change, and it calls for transnational regulatory bodies.

That is exactly what Bengio proposed: that developers demonstrate to independent experts that their systems are safe to train and safe to deploy, before doing either; that security incidents be reported; that frontier systems be licensed, like medicine, aviation or nuclear energy. “Decisions that affect us all should be made by us all.”

We can debate the timeline, or how likely the worst case is. But a parent who catches their child lying doesn't wonder whether the teenage years will be difficult; they start preparing. The time to set the rules is before adolescence, not during. In that sense, Yoshua Bengio's call is not alarmist. It comes right on time.

At our scale, it's also a design principle: our agents only act within a scope that you define, and never directly inside your systems. Want to discuss it? Let's talk.

— Frédéric Brabant

All articles