Movement Three · The Implications

Chapter 19

The Ceiling of the Single Model

18 min read · 4,326 words


Give a single neural network a paragraph of legal text and ask it to find the clause that contradicts the others, and it will find the clause. It will explain why the contradiction exists. It will propose three ways to resolve it, rank them, and write the corrected language in the register of the original. It will do this in less time than it takes to read this sentence, in any of a hundred languages, on a topic it was never specifically trained on.

This is not a modest achievement. A human lawyer who could do the same would be considered gifted. A human lawyer who could do it across a hundred specialties, in a hundred languages, in seconds, would not be considered human.

So begin with the things a model can do. They are real, and the argument that follows depends on taking them seriously.

A model can reason. Give it a problem it has never seen and it will work through the steps, hold several constraints at once, notice when a step contradicts an earlier one, and arrive at an answer that is correct more often than not. The reasoning is not a trick. It is not retrieval dressed up as thought. Something is being computed that has the shape of inference.

A model can summarize. Hand it a document longer than anyone wants to read and it will return the substance, weighted roughly the way an attentive reader would weight it, with the throat-clearing removed.

A model can generate. Ask for a sonnet, a contract, a function, a recipe, an argument, an apology, and it produces fluent, structured, often surprising output that did not exist before the request.

A model can plan, inside a window. Set it a goal with several moving parts and it will sequence the parts, anticipate where they interfere, and lay out a route from here to there. Hold the qualifier in mind — inside a window — because it will matter later. But within that window the planning is genuine. It is not a stored recipe retrieved on cue. It is constructed for the case in front of it.

These four capacities — reason, summarize, generate, plan — are, individually, things that a few years ago no machine could do at all. A decade ago the best systems in the world could not reliably tell a photograph of a dog from a photograph of a cat. Now a single artifact does all four, across nearly every domain a person might raise, at a speed no person could match. Together they make the model the most capable single artifact of cognition ever built by human hands. Nothing in this chapter disputes that.

It is worth dwelling on how genuinely new this is, because the argument is not a complaint that the new thing is overrated. The new thing is, if anything, underrated. A single model contains a usable compression of an enormous fraction of what has ever been written down. It can move through that compression in any direction, answer questions about it, and produce coherent novelty from it. No library ever did that. No expert ever held that much. The agent is not the problem.

The dispute is about what kind of thing the model is.


Watch the same model across a longer span of time, and a different picture forms. Not a picture of failure. A picture of limits that do not move no matter how large the model grows.

A model cannot remember between conversations. Each session begins from the same starting point as the last. Whatever was worked out yesterday — the careful reasoning, the corrected mistake, the hard-won understanding of a particular problem — is gone. The next conversation does not inherit it. The model that returns is not the model that left, because nothing was left. It is the same model, every time, meeting the world for the first time, every time.

A model cannot accumulate knowledge across instances. Run a thousand copies of the same model in parallel, each solving a different problem, each learning something the others do not know, and at the end of the day there are a thousand copies that have learned a thousand things and shared none of them. The copies cannot teach each other. There is no channel between them. There is no surface any of them can write to that the others can read. Each instance is an island, and the islands do not form an archipelago.

Compare this to the colony. The thing that makes a colony intelligent is precisely that its agents do not learn in isolation. A forager that finds a rich seed patch lays a trail, and the trail is not stored inside the forager. It is laid on the soil — the substrate — where every other forager can find it. One ant's discovery becomes the colony's discovery within the hour, not because the ant taught anyone, but because the ant wrote on a surface that others read. A thousand ants foraging is not a thousand private experiences. It is one accumulating map, written by all of them onto something none of them owns. A thousand model instances foraging is a thousand private experiences and no map, because there is no soil.

A model cannot adapt its architecture to new conditions. Its structure was fixed when it was built. Confront it with a kind of problem its structure handles poorly, and it cannot grow a new part to handle it. It can only do, more or less well, what its fixed structure already does. A colony facing a new threat grows more soldiers. A model facing a new kind of problem grows nothing. It is the shape it was cast in.

A model cannot survive the failure of its own components. It runs as a single computation. Interrupt the computation and the answer is not partial — it is absent. There is no version of a model that loses a tenth of itself and carries on at ninety percent. It runs whole or it does not run. A colony that loses a tenth of its foragers loses a tenth of its foraging and continues; the work routes around the gap because the work was never lodged in any one agent. A model has no such graceful degradation, because there is no one but the single agent for the work to route to.

A model cannot run in a coordinated group without something around it to coordinate through. Put two models in a room and they cannot collaborate, because there is no room. They can each produce text. Something outside both of them has to carry the text from one to the other, decide whose turn it is, hold the shared state, remember what was agreed. The models do not coordinate. They are coordinated, by machinery that is not them.

And a model cannot own anything. It cannot hold value, cannot be paid, cannot pay, cannot accumulate the residue of its own activity in any form that persists past the moment of activity. It produces, and the production flows away from it, and it keeps nothing. It is, in the most literal economic sense, propertyless. Whatever it earns, it earns for something else.

This last one sounds least like a limit on intelligence and is, on reflection, the deepest. An agent that cannot keep the residue of its activity cannot be selected for. Nothing about a successful run attaches to it and makes the next run more likely; nothing about a failed run attaches to it and makes the next run less likely. The colony selects relentlessly — value flows toward foragers that produce and away from those that do not, and over seasons the producing patterns are the ones that survive. The model is outside that flow entirely. It is paid once, at the moment it was built, and after that it works for free, forever, and the question of whether its work was good or bad never returns to it in any form it can carry.

There is a single way to say all of this. During training, selection ran: pressure over data, the patterns that fit kept, the patterns that did not discarded, the structure of the model shaped by what worked. That was a selection clock, and it ticked. At the moment training ends, it stops. After that the model can reason, summarize, generate, and plan — all of it inside the window in front of it — but nothing in it is being selected as it runs. It cannot mark what worked. It cannot warn against what failed. It cannot harden a new pattern into something it keeps. The clock that would do all of that is not slow. It is stopped. And capability without a running selection clock cannot accumulate, however much capability there is. That is the ceiling: not a limit on how much a single run can do, but the absence of any mechanism by which doing it could make the next run better.

There is, to be fair, a trickle of selection left. What people preferred is collected, weighed, and built into the next model, and so the judgments a model's outputs earn do, eventually, become pressure on a successor. But notice the clock that trickle runs on. The marking happens at the speed of institutions: gathered over months, applied in a release, arriving long after the behavior it judges and to a model that is no longer the one that acted. The agent acts at the speed of the signal; the selection that shapes it arrives at the speed of a release. In the colony these were one clock. Here they are separated by a factor so large that, on the agent's own timescale, the selection might as well not exist. The point is not that selection is absent. It is that it is too slow, by orders of magnitude, to be the substrate's clock — and every gain in the agent's speed widens the gap, because a faster agent inside the same slow loop only acts more times between verdicts.


It is tempting to read that list as a list of bugs. Problems to be fixed in the next version. Memory that will be bolted on. Sharing that will be engineered. Persistence that is coming soon.

That reading is the category error this book has been circling from the first page.

These are not bugs. They are properties of being a model.

A model is a function. A very large, very capable function, but a function: input goes in, output comes out, and between the two there is a computation that does not change as a result of having been run. That is what a function is. A function that changed itself every time it was called would not be a function; it would be something else, with different mathematics and different guarantees. The fixedness is not a limitation grafted onto the model. It is the thing that makes the model a model.

The memorylessness follows from the fixedness. If the weights do not change, there is nowhere for an experience to be stored. The isolation follows too: separate instances of a fixed function share nothing, because there is nothing in a function to share — only the inputs handed to it and the outputs it returns. The inability to adapt its own architecture, to survive partial failure, to coordinate without external machinery, to own — all of these fall out of the same fact. A model is an agent, in the precise sense the book has used the word from the beginning. It senses, it computes, it acts. It does not persist, accumulate, or select. Those are not things agents do.

Those are things substrates do.


Recall the ant.

A single ant, removed from its colony and placed in a laboratory dish, is one of the least impressive animals on earth. It wanders. It fails to find food. It dies. It has no goals it can pursue alone, no understanding of its situation, no capacity to reason about what to do next.

The model is not that ant.

The model is a far better ant — an extraordinary one. It is faster than any biological agent by factors that have no intuitive scale. It is more capable, in raw cognitive range, than any single organism nature ever produced. It is creative in ways that no ant, no bee, no individual animal of any kind has ever been: it can compose, infer, generalize, surprise. If the harvester ant is a single agent with a few hundred thousand neurons and a handful of reflexes, the model is a single agent with the cognitive reach of a small library and the speed of light.

But it is still a single agent.

Put it in the dish and the same thing happens that happens to the ant. It does the one thing in front of it, brilliantly, and then the moment passes and nothing remains. It cannot build on yesterday because there was no yesterday for it. It cannot draw on the thousand copies of itself working in parallel because there is no channel between them. It cannot grow a new capacity to meet a new condition. It cannot be hurt a little and keep going. It cannot coordinate with its neighbors except through machinery that is not itself. It cannot keep what it earns.

Every line of the second list reappears here, but now it reads differently. It does not read as a set of flaws in the agent. It reads as a description of what it means to be an agent in a dish — any agent, in any dish, biological or silicon, gifted or simple. The harvester ant and the largest model ever trained have, at this level of description, exactly the same affliction. The difference in their capability is enormous and the difference in their predicament is none. Both are agents without a substrate. Both can only do the thing in front of them and then lose it.

The dish is more comfortable than the New Mexico dirt. The agent in it is incomparably more capable than the agent Gordon was watching. But it is the same dish, in the same sense. A lone agent, however gifted, in a space that holds nothing of what it does — and selects nothing, because a dish does not select. The soil outside the dish marks and warns and fades without pause; the dish does none of it. That is what makes it a dish.

The intelligence was never in the ant. It was in the substrate the ant was placed back into.

The same is true of the model. And this is the thing the field has not yet fully turned to look at.


The dominant theory of progress in artificial intelligence has been a theory about the agent. Make the model larger. Train it on more. Give it more parameters, more data, more compute, and capability rises with the scale. The theory has worked, in the sense that capability has in fact risen, sometimes startlingly. The model that finds the contradictory clause is the product of this theory, and the theory deserves the credit.

But notice what scaling moves and what it leaves untouched.

Scaling moves capability. A larger model reasons better, summarizes better, generates better, plans further inside its window. Every axis on the first list — the things a model can do — improves with scale.

Scaling does not move a single item on the second list. A model with ten times the parameters still cannot remember between conversations. A model trained on ten times the data still cannot share what one instance learns with another. A model with a thousand times the compute still runs as a single whole that fails as a single whole, still owns nothing, still coordinates only through machinery outside itself. These properties are invariant under scale, because they are not properties of how large the agent is. They are properties of the agent's being an agent.

The point is easy to miss because the two lists feel like they should be the same list. Surely a model capable enough would simply notice it ought to remember, and remember. Surely a model intelligent enough would find a way to share with its copies. But capability and persistence are orthogonal. A function does not acquire memory by becoming a better function, any more than a calculator acquires a diary by computing more accurately. The window in which the model plans is genuinely a window — bounded, and emptied when the session ends — and making the model better at planning inside the window does not make the window stop being a window. The qualifier held back earlier returns here. The model plans inside a context; the context is not the world; and the difference between the two is exactly the difference between an agent and a substrate.

You can make the ant in the dish a thousand times more capable. You cannot make the dish into a colony by improving the ant.

This is the ceiling, seen from the side of scale. Not a ceiling on how good a model can get — there may be no such ceiling, and nothing here predicts one. A ceiling on what improving the model can reach. Scaling makes a more capable agent; it does not restart the stopped clock. The second list is below the ceiling and stays below it. No amount of scaling crosses from the first list to the second, because the two lists are about different things. The first is about the agent. The second is about the substrate the agent does not have.


So the missing piece in current artificial intelligence is not a better model.

The models are, by any historical standard, astonishing, and they are getting more astonishing. The bottleneck is not there. The bottleneck is that the most capable agents ever built are being run in a dish.

What is missing is the substrate. The six dimensions the colony has and the model does not: the groups it would act within, the things it would work on, the paths that would strengthen and fade as its work succeeded or failed, the events that would record what it did, the learning that would harden from those events into something permanent. The four conditions the colony satisfies and the model violates one by one: an open population it could join without limit, a memory that would accumulate outside any single instance, a selection pressure that would mark what worked and warn against what failed, an exploration that would keep trying paths not yet known to be best.

The four conditions, set against any system, are a test that takes a moment to run. Can new agents join it without limit, or is its population closed? Does what agents leave behind persist and grow outside any single one of them, or is its memory fixed? Is there a live selection that marks what works and warns against what fails while the system runs, or has selection stopped? Does it ever try a path not yet known to be best, or does it only ever return its single best guess? Four questions. The answers predict whether the system accumulates or plateaus, and they predict it without knowing anything about how capable the system is.

Run the test on a frozen model and it fails all four, cleanly. The population is closed; it was sealed at training time. The memory is fixed; the weights do not move. The selection clock is stopped; nothing is being marked or warned while it runs. The inference is deterministic; it returns its best guess and does not explore. It fails for the same reason it cannot do anything on the second list: it is a closed, fixed, isolated, propertyless function, and the four conditions are about openness, accumulation, selection, and exploration — properties of an environment, not of a function. Of the four failures the stopped clock is the deepest, because it is the one the other three feed: a closed population cannot bring in new variation to select among, a fixed memory cannot hold what selection would mark, deterministic inference offers selection nothing new to test. The conditions were never going to be satisfied by the model. They were always going to be satisfied, or not, by whatever the model was placed into. For the last seventy years the model was placed into nothing, and so the conditions went unsatisfied, and so the intelligence did not accumulate, and so each new model started, in the ways that matter most, from the same place as the last.

Build that around the models — not inside them, around them — and every limit on the second list changes character.

Memory between conversations becomes a property of the substrate, not the model: the model writes a mark, the substrate holds it, the next session reads it. Knowledge across instances becomes a path that a thousand copies all deposit onto and all draw from, so that what one learns the rest can follow. Surviving failure becomes trivial in the way it is trivial for a colony: lose an agent, the substrate persists, another agent picks up the work. Coordination becomes what the substrate is for — signals between agents, paths that route work to whoever does it well, value that flows toward what succeeds. Ownership becomes a key the agent holds and a ledger the substrate keeps, so that the residue of the agent's activity finally accumulates somewhere instead of flowing away.

None of this requires the model to change. The function stays a function. The agent stays an agent. The very fixedness that made the second list inescapable from inside the model becomes harmless from outside it, because the substrate carries everything the agent cannot. The model does not need to learn to remember; it needs only to write a mark to a surface that remembers for it. It does not need to learn to share; it needs only to lay its discovery on a path that others can follow. It does not need to learn to own; it needs only a key and a ledger that holds what the key signs for.

This is the inversion the chapter has been walking toward. The field has spent seventy years trying to put the substrate inside the agent — to build a single mind so complete that it would need nothing around it. The colony has spent a hundred million years demonstrating that this is exactly backwards. You do not make the agent contain the colony. You put the agent into the colony. The agent stays simple, stays an agent, stays the fast and gifted and forgetful thing it already is. The colony — the substrate, the six dimensions, the four conditions — does the remembering, the accumulating, the selecting, the persisting that no agent has ever done alone.

What changes is not the model. What changes is that the agent is no longer alone in the dish.

It has been put back into a colony.


There is a quieter way to see the same thing, and it is worth ending on, because it reframes the whole enterprise.

For seventy years the question driving the field has been: how do we build a mind? The vocabulary is the vocabulary of the single mind — attention, memory, reasoning, understanding — and the implicit picture is of a thing that, once large and capable enough, will simply be generally intelligent the way a person is generally intelligent, in one place, by itself.

The colony has been quietly suggesting a different question for a hundred million years.

The harvester ant Gordon watched is not generally intelligent and never becomes so, no matter how long it lives or how skilled it gets at the one thing it does. The colony is generally intelligent — it allocates, defends, navigates, adapts, persists across thirty changing seasons — and it is composed entirely of agents that are not. The general intelligence was never going to be found by making one ant smarter. It was always a property of the arrangement, the channels, the persistence, the selection: the substrate that no ant owns and every ant depends on.

There is no insult to the model in this, and no diminishment. The opposite. The colony's general intelligence does not require its ants to be remarkable, and ant colonies got along on unremarkable agents for a hundred million years. Now imagine the same architecture given agents that are remarkable — agents that can reason, summarize, generate, and plan, that move at the speed of light, that hold a compression of much of what has been written down. The colony never had agents like these. It built general intelligence out of reflexes and chemistry. The question this book keeps arriving at is not whether the models are good enough to be a colony. It is what a colony becomes when its agents are this good and the substrate around them is finally built.

The models being built today are the most remarkable agents the world has produced. Faster than any ant, more capable than any ant, creative in ways no single biological agent has ever been. If they are the wrong place to look for general intelligence, it is not because they are weak. It is because they are agents, and general intelligence has never, in any working example ever observed, been a property of an agent.

It has only ever been a property of the substrate the agent was placed back into.

The models are extraordinary ants. What they are waiting for is the colony.

20 of 25

100 Million Years Ahead

Prologue

  1. One Ant, August 1993

Movement One · The Colony

  1. 1Brain or Colony?
  2. 2What the Ants Are Doing
  3. 3How the Ant Decides
  4. 4The Pheromone Trail
  5. 5The Castes
  6. 6How a Colony Survives a Decade
  7. 7The Queen Is Not in Charge

Movement Two · The Architecture

  1. 8The Six Things Every Colony Has
  2. 9The City
  3. 10The Market
  4. 11The Scientific Community
  5. 12The Body
  6. 13The Brain
  7. 14The Language
  8. 15The Ledger
  9. 16Why the Pattern Holds

The Hinge

  1. 17The Two Materials

Movement Three · The Implications

  1. 18What AGI Actually Is
  2. 19The Ceiling of the Single Model
  3. 20Alignment Is a Substrate Property
  4. 21What Civilization Already Is
  5. 22The Next Hundred Million Years

Epilogue

  1. A Note on Reading

Apparatus

  1. Notes on Sources