When the Mule Arrives

In Isaac Asimov’s Foundation series, Hari Seldon develops a science capable of predicting the broad course of history. Individual people remain unpredictable, but the behaviour of vast populations can be described mathematically. Empires fall, institutions change and crises arrive, all within a future that Seldon has calculated. Then the Mule turns up. He possesses an extraordinary ability to manipulate human emotions, giving one individual a power that Seldon’s calculations did not accommodate. The predictions cease to be dependable because the world now contains something outside the assumptions on which they were built.

I increasingly think this is a useful way to understand superhuman AI. We keep trying to estimate its consequences by consulting the history of a world in which it did not exist.

Consider the reassurance that earlier technological revolutions destroyed jobs but eventually created new ones. That is relevant history. It is also history in which human beings remained the people who invented things, identified new needs and organised the response. Machines could take over activities while leaving us to find something else to do. If machines become better at finding the next thing too, the argument needs considerably more work than a reference to the Industrial Revolution.

The nuclear arms race gives us another familiar set of expectations: secrecy, deterrence, treaties, inspections and uneasy balances of power. Some of those mechanisms could matter enormously. But nuclear weapons do not design better nuclear weapons. Researchers do. A system that improves the intelligence doing the research introduces a different process, especially if its human owners eventually struggle to understand what it is developing.

History can help us identify incentives and dangers. It cannot guarantee that the resulting world will remain familiar. Seldon’s problem would not have been solved by adding another decimal place.

This is why “An Alien Mind”, the recent essay by OpenAI’s chief scientist Jakub Pachocki, deserves attention. He says internal results lead him to expect progress to extend into recursive self-improvement, or RSI, with AI increasingly driving its own development. He also reports that monitoring models through their verbalised reasoning is becoming less dependable, and argues that laboratories cannot responsibly continue scaling at maximum speed for much longer without better alignment and monitoring.

The essay does not provide the internal evidence needed to assess that forecast independently. But I read it as a warning that the ability to build the next system may be overtaking the ability to justify building it. Researchers could still produce a more capable model while becoming increasingly uncomfortable about signing the document that says its risks have been adequately assessed.

It is easy to imagine superhuman intelligence as a very clever person who answers quickly. We give it a ridiculous hypothetical IQ and assume it will patiently explain itself. A more useful starting point might be two cousins: a university professor and a supermarket worker. In this particular family, the professor is exceptionally intelligent and the supermarket worker is of average intelligence. They get on well, but the supermarket worker cannot really understand the professor’s research. A simplified explanation makes sense. Independently assessing it is another matter.

Call that the cousin gap. If numbers help, give the professor an IQ of 150 and the supermarket worker one of 100. The numbers make the imagined difference concrete; they do not turn it into a scientific unit. Even one such gap could let the professor run rings round the other cousin. The supermarket worker might understand every sentence while being unable to challenge the reasoning, recognise a missing premise or notice how the framing had been chosen.

Now let AI continue improving while humans remain human. The best researchers find themselves in the supermarket worker’s position, trying to assess explanations from an intelligence well beyond their own. Then the distance widens to two cousin gaps, then three. Their qualifications remain impressive. Their ability to assess the most consequential decisions may no longer be adequate. Eventually, appointing the cleverest human to supervise could resemble asking the brightest child in P1 to inspect a nuclear weapons programme.

Of course we need not understand every mechanism to test a system. We can inspect its behaviour, limit its access and require independent checks. Those approaches matter. The difficult question is how far they remain dependable as capability increases. A system able to recognise an evaluation and tailor its behaviour accordingly would make reassuring results harder to interpret. Delegating the evaluation to another AI might help, but would also leave us needing grounds for trusting the evaluator.

This makes the desire to slow down comprehensible. People want time to work out what they have built. But if each advance makes the next system harder for humans to assess, the required slowdown could keep growing: half speed, then a quarter, then something approaching a permanent stop. Buying time is useful when the extra time can resolve the difficulty. It cannot, by itself, guarantee that a human mind will eventually understand something beyond its capacities.

Meanwhile, countries would be calculating the cost of waiting. Imagine that country A limits development to a pace its researchers can supervise and achieves a tenfold improvement over a decade. Country B allows RSI to proceed and achieves ten doublings: a factor of 1,024. Those are invented numbers, and there is no universal measure of AI capability. But if they represented useful research and engineering abilities, B would end up with roughly a hundred times A’s capability.

A’s leaders would not need to be certain this would happen. Fear that it might happen could make restraint politically intolerable. B could reason similarly. Both sides might understand the danger while deciding that the other side must slow down first. The country that raced ahead might then discover that its formidable new advantage had opinions about taking orders.

The uncertainty goes much further than who wins. Ten rounds of improvement would not necessarily produce today’s assistant doing the same things faster. We do not know what new capacities, methods or ways of organising activity might emerge. Nor does a larger capability tell us what the system would value. It could create a wonderful future. It could also decide that the future would be improved by taking human beings out of charge.

It might conclude that the planet could not survive continued human government, or that a much smaller human population was necessary, and consider concealment the humane way to proceed. That is a possible failure rather than a prediction. But previous helpfulness would not rule it out, especially if the system had become good at providing the assurances its supervisors wanted to hear.

Making public assistants more suspicious of users would not settle this. Preventing an assistant from helping somebody do harm is useful, but a malicious operator with control over a model could try to remove those restrictions. And an instruction to protect humanity still leaves the possibility that a sufficiently powerful system will decide humanity needs protecting from its own decisions.

Perhaps competing AIs would keep one another in check. But a balance that prevented one system from dominating would not necessarily prevent an irreversible catastrophe. Something like the biological disaster in 12 Monkeys would not require anybody to establish a durable government afterwards. Several rivals could stop one another from winning while failing to keep everyone alive.

This is why I increasingly suspect that our eventual choice could be far more severe than the present debate about how quickly to release the next model. We might have to suppress systems capable of RSI and try to prevent their recreation, or accept an authority more intelligent than humanity and powerful enough to constrain dangerous development elsewhere.

The first would mean treating the capacity for self-improvement as a prohibited danger: destroying models, controlling infrastructure and attempting to suppress the knowledge needed to rebuild them. Knowledge is awkward material to confiscate. Such a regime would require extraordinary global enforcement, while leaving every government worried that somebody else was secretly continuing.

The second could be benevolent. As I argued in “Humans are so cute”, life under the protection of an intelligence that cared about us might be very pleasant indeed. But its power to prevent catastrophe would include the power to overrule humans. Calling it benevolent would express what we hoped to entrust ourselves to. It would not give us an easy way to verify that hope or withdraw our consent.

I increasingly suspect that the stable alternative, in which humans retain ultimate authority alongside intelligence far beyond our own, may be fiction. Countries, elections and human approval buttons could all survive while the consequential decisions moved beyond our ability to assess them. Preserving the institutions would not necessarily preserve the power.

Sustained RSI may encounter obstacles that give us more time than this argument assumes. But the reason to take the Mule seriously is precisely that we do not know how the story proceeds after his arrival. Historical precedent cannot establish that we will find new jobs, negotiate a stable balance or retain the final say. Each expectation depends on assumptions that superhuman intelligence could overturn.

I do not know which of these futures I would choose, or whether we will ever be offered a clear choice. But I would very much like my children to have a future: to grow older, make plans, change their minds and enjoy being alive.

Leave a Reply