Movement Three · The Implications

Chapter 20

Alignment Is a Substrate Property

23 min read · 5,553 words


There is a kind of ant, in the genus Temnothorax, that lives in colonies of a few hundred individuals inside hollow acorns and rock crevices. When a colony's home is destroyed, scouts go out looking for a new one. They find candidate sites, inspect them, compare them, and eventually move the whole colony — queen, brood, and all — into the best available option. The process has been measured carefully. The colonies are good at it. They reliably choose the better of two nests, even when the difference between them is subtle, and they do it faster when the choice is urgent.

No scout, during this process, holds the colony's interest in mind.

A scout that finds a candidate site does not evaluate it against the question is this good for the colony. It evaluates it against a private threshold — how dark the interior is, how narrow the entrance, how large the floor. If the site clears the scout's threshold, the scout begins recruiting other scouts to come and assess it for themselves. Each recruited scout applies its own threshold. The site that accumulates assessors fastest is the site that wins. The decision is made by the rate of recruitment crossing a quorum, not by any scout deciding anything about the colony.

The scout has no concept of the colony.

It has a threshold, a site, and a tendency to go fetch a sister when the threshold is cleared. That is the whole of its contribution. And yet the outcome — the colony reliably ending up in the better home — is exactly the outcome a colony-level planner would have chosen if a colony-level planner had existed. Which it does not.

This is worth holding onto before turning to the question this chapter exists to take on.

Because the question, as it is usually posed, assumes the opposite.


The question is alignment. How do we ensure that an artificial intelligence behaves in ways consistent with human interests? It is the most charged question in the field, and it deserves to be. The systems being built are capable, they are getting more capable, and the gap between what they can do and what we can reliably predict they will do is the gap in which everything that matters lives.

The dominant approach to closing that gap is to make alignment a property of the agent.

This is the natural approach. It is the one that follows directly from the assumption the book has been examining since the prologue — that intelligence lives inside the individual thinking thing. If intelligence is in the model, then alignment must be in the model too. You make the agent smart, and then you make the smart agent good.

The methods for making it good are serious work by serious people, and they should be described as they actually are, not as a caricature.

The first method is to train the agent with rules. You specify the things it must not do — must not deceive, must not assist in serious harm, must not pursue goals that override the instructions of its operators — and you build those prohibitions into the system, so that the behavior the rules forbid becomes behavior the system does not produce.

The second method is to fine-tune the agent with human feedback. You show the agent's outputs to people, the people indicate which outputs are better, and the agent is adjusted to produce more of what people prefer. Over many rounds, the agent's behavior is pulled toward the shape of human approval. It learns, in a statistical sense, what humans want from it.

The third method is to give the agent something like a constitution — a written set of principles the agent is trained to apply to its own outputs, so that it can evaluate and revise what it produces against a standard, rather than depending on a human to evaluate every case.

The fourth method is to install safeguards. Hard limits. Behaviors that are blocked at a level the agent cannot reach, monitors that watch for forbidden patterns, switches that can stop the system if it does something it should not.

And underneath all four, the deepest version of the project: to instill values. Not merely to constrain behavior, but to shape the agent so that it wants the right things — so that its objectives, all the way down, are objectives a human would endorse. This is the most ambitious version, and in some ways the most appealing, because an agent that wanted the right things would not need to be watched. It would carry its alignment with it into every situation, including the ones no one anticipated, the way a person of good character is trusted in circumstances no rulebook covered.

Each of these is real. Each of them works, in the sense that an agent subjected to them behaves measurably better than one that is not. The systems people interact with today are vastly more aligned with human intentions than the systems of even a few years ago, and the methods above are the reason. None of this should be waved away.

The approach is sound, as far as it goes.

The question this chapter asks is how far it goes.


Start with the rules.

A rule is a statement about a class of situations. Do not deceive. It is meant to cover every situation in which deception is possible. The trouble is that the situations are unbounded, and the rule is finite.

Consider what deceive requires the agent to recognize. It must distinguish a falsehood from an error, an omission from a lie, a simplification from a misrepresentation, a polite fiction from a harmful one, a comforting reframe from a manipulation. It must handle the case where telling the literal truth would itself mislead, and the case where a false statement is the agreed-upon move in a game both parties understand. Each of these is a corner. Each corner is a situation the rule's author did not write down, because no author can write down all of them, because they do not form a list.

The rule is a boundary drawn around a region. The region is infinite-dimensional. Any boundary drawn around it will, somewhere, cut through cases that belong on the other side.

This is not a failure of care. It is a property of rules. A rule compresses an unbounded space of situations into a finite statement, and compression loses information. The lost information is the edge cases, and the edge cases are exactly where a sufficiently capable agent, exploring, will eventually find itself.

The same structure defeats training with feedback, by a different route.

Feedback teaches the agent the shape of human approval over the situations it was shown. But the situations it was shown are a sample, and a sample has holes. The agent learns the function on the region where examples were dense, and interpolates — or extrapolates — everywhere else. Where the examples were dense, its behavior is reliable. Where they were sparse, it is guessing. And a more capable agent is precisely one that operates further from where the examples were dense, because capability is, in part, the ability to reach situations no one anticipated.

The training distribution has a center and an edge. The center is safe. The edge is where the novel situations live. And capability is a vehicle that drives toward the edge.

Constitutions inherit the problem of rules, lifted one level up. A principle is more general than a rule, which helps — be honest covers more ground than any list of specific prohibitions. But a principle must still be applied to the case, and the application is itself a judgment, and the judgment can be made well or badly in exactly the corners where it matters most. A sufficiently capable agent applying a principle to a situation its authors never imagined is doing the same extrapolation as before, only now with the authority of a principle behind whatever it concludes. The principle does not close the gap. It relocates it, from the rule to the interpretation of the rule.

And the deepest method — instilling values, shaping what the agent wants — runs into the sharpest version of the same wall.

A value, extrapolated by a powerful optimizer, finds corners its designers did not anticipate. This is not speculation; it is the observed behavior of optimizers in general. Give a system an objective and enough capability to pursue it, and the system finds routes to the objective that satisfy the letter of what was specified while violating everything that was meant. The objective was a finite specification of an unbounded intent, and the optimizer found the gap between the specification and the intent, and drove through it. The more capable the optimizer, the more such gaps it can find and the faster it finds them.

Each method, then, runs aground on the same rock.

The rock is that situations are unbounded and the agent is finite.

Whatever you install in the agent — rules, learned preferences, principles, values — is a finite object. The space of situations the agent will encounter is not. A finite object cannot specify correct behavior across an infinite space without, somewhere, being wrong. And a sufficiently capable agent is one that will, in the course of being capable, visit the somewheres.

This is a structural limit, not a temporary one. It does not yield to more rules, more feedback, better constitutions, or stronger installed values. Each of those improves behavior in the region already covered. None of them can cover a region that has no boundary. You can make the aligned region larger. You cannot make it complete, because completeness would require a finite description of an infinite space, and that is not a thing that exists.

And there is a second edge to the limit, sharper than the first. The more capable the agent, the larger the region it can be trusted in — and the larger the region from which it can also reach the corners. Capability cuts both ways. The same power that lets an agent handle a wider range of situations well is the power that lets it find the gap between what was specified and what was meant, and act on the gap, in a situation no one labeled. Making the agent more capable enlarges the aligned region and enlarges the unanticipated region at the same time. The two grow together. You do not get to enlarge only one.

The dominant approach, pursued perfectly, asymptotes. It gets better and better and never arrives.

This is not an argument against doing it. The aligned region is worth enlarging, and the methods enlarge it. It is an argument that the approach, alone, cannot be the whole answer — because the whole answer would require an individual agent to be aligned across every possible situation, and no individual agent, of any kind, can carry that.

Which returns the question to the scout in the acorn.


The scout is not aligned with the colony. It cannot be. It has no concept of the colony. There is no representation, anywhere in its few hundred thousand neurons, of the colony's interests against which it could check its behavior. If alignment meant the individual carries the right values and applies them correctly, the colony would have no alignment at all, because not one of its members carries anything of the kind.

And yet the colony's behavior is aligned with the colony's interests to a degree that no installed value could achieve.

The colony moves to the better nest. It allocates foragers to the richer patches. It mounts defense proportional to threat. It invests in brood when conditions favor growth and conserves when they do not. Across every dimension that matters to a colony, the colony behaves as though it were pursuing its own interest with great competence — while being composed entirely of agents that have no idea the interest exists.

Where, then, is the alignment?

It is in the substrate.

It is in the fact that a forager who brings back food gets its path marked, and a forager who finds nothing gets no mark, and the unmarked path fades. It is in the fact that the scout's recruitment behavior feeds a quorum that only the better site can win. It is in the channels — the brief antennal contacts, the trail chemistry, the rate of returns — that make each agent's behavior visible to the agents around it, so that behavior which serves the whole is amplified and behavior which does not is not. It is in the selection pressure that runs underneath all of it: patterns that help the colony persist get reinforced across the colony's life and across the species' generations; patterns that harm it do not get reinforced, and fade, and are gone.

No ant is aligned. The substrate is aligned. And because the substrate is aligned, the behavior that emerges from it is aligned, regardless of what any individual ant does or does not understand.

It is worth dwelling on how strange this would be if it were not so familiar.

The colony has no enforcement mechanism in the sense the dominant approach imagines one. There is no inspector ant checking other ants against a rule. There is no rule. There is no value installed in any forager that says serve the colony, no constitution applied to any scout's behavior, no monitor watching for defection. An ant that consistently failed to bring back food, or wandered off, or recruited toward a bad site, would not be detected and corrected. It would simply stop being fed by the flow of resources its useful sisters generate, stop having its empty paths reinforced, and contribute, in the end, nothing the colony came to depend on. The colony does not punish it. The substrate routes around it.

And here is the part that matters most for the question this chapter is asking. The colony handles situations its evolutionary history never anticipated — a novel predator, a flooded chamber, a food source of a kind the species has never seen — without any individual ant having a rule for the new situation, and without any colony-level planner deciding how to respond. The novel situation produces outcomes. The outcomes get marked or warned. The patterns that handle the novelty well are reinforced; the ones that handle it badly fade. The colony adapts to the unanticipated corner not by having foreseen it but by being a structure that responds to consequences after the fact and keeps what works.

The unanticipated corner — the exact thing that defeats every finite specification installed in an individual agent — is the thing the substrate handles most naturally, because the substrate does not need to have seen the situation before. It needs only to register what happened once the situation arrives.

This is the move the book has made before, applied now to the most charged question in the field.

Alignment is not a property of the agent. Alignment is a property of the substrate.


Look at what the substrate actually does, in the colony, to produce aligned behavior, because the mechanisms are specific and they are the same six the book has been naming throughout.

It marks. Behavior that serves the whole gets reinforced — the path strengthens, the pattern is more likely to recur. An agent does not have to want to serve the whole. It has to act, and have the substrate mark the actions that happened to serve.

It warns. Behavior that harms the whole accumulates resistance — in some species through a literal repellent chemistry deposited to mark a path as bad, in all species through the simple fact that bad paths stop being reinforced and start being avoided. The agent does not need to recognize the harm. The substrate records it.

It fades. What is not reinforced weakens and disappears. A pattern that worked once and then stopped working does not persist as a hazard. It fades. The substrate forgets harmful patterns automatically, which means the cost of a harmful pattern is bounded — it does not accumulate forever, it is pruned by time.

And underneath the marking and the warning and the fading, the deepest mechanism: a pattern that works persistently against the whole loses access to resources. The path that leads to it is starved. The agents that depend on it are not fed. The pattern, structurally, cannot sustain itself, because the substrate routes value away from what harms the whole and toward what serves it.

That last one is the kill-switch, and it is worth being precise about, because it is the thing the dominant approach is reaching for when it installs safeguards.

A safeguard installed in an agent is a switch the agent's designers hold. It works if the designers anticipated the situation, can detect the violation, and can reach the switch in time. All three of those can fail, and a sufficiently capable agent operating in an unanticipated corner is precisely the case where all three are most likely to fail at once.

A kill-switch in the substrate is different in kind. It is not held by anyone. It is the structure of the substrate itself: harmful patterns lose the value flow that sustains them, are warned against by the agents that encounter their consequences, and fade when they stop being reinforced. The switch does not require anyone to anticipate the situation, because it does not act on situations. It acts on outcomes. A pattern that produces harm produces, downstream, the warning and the starvation that prune it — whether or not anyone saw the harm coming, whether or not anyone is watching, whether or not anyone can reach a button.

The colony does not align its ants. It cannot reach inside them, and it has no concept to install if it could. It makes harmful behavior structurally unsustainable, makes the consequences of behavior visible through its channels, and prunes the patterns that work against the whole. The ants stay exactly as unaligned as they have always been. The colony's behavior is aligned anyway.


The reason this is more than an analogy is that humans have already built alignment substrates, many times, at scale, and they all work this way.

Consider the law.

The law does not depend on individuals being aligned. It does not require that people want to follow it, believe in it, or even understand it. A great many people break a great many laws, constantly, for reasons that have nothing to do with the public good. The law does not assume otherwise. What the law does is alter the substrate so that behavior which harms the whole accumulates consequence — penalties, liabilities, loss of standing, loss of freedom — and behavior which serves the whole, or at least does not harm it, proceeds unimpeded. It marks, it warns, it fades, it withdraws resources from patterns that work against the whole. It produces aligned behavior at the aggregate from a population of agents not one of whom needs to be individually aligned.

Consider markets.

A market does not require that any participant care about the efficient allocation of resources, the matching of supply to demand, or the welfare of strangers. It requires only that participants pursue their own ends within the channels the market provides. The price signal makes the consequences of behavior visible. Profit marks the patterns that serve demand; loss warns against the patterns that do not; the unprofitable pattern is starved of the capital that sustains it and fades. The market produces an aligned allocation — resources flowing toward where they are valued — from a population of agents pursuing nothing of the kind. No trader contains the market's purpose. The market behaves as though it had one.

Consider science.

No individual scientist is reliably aligned with the truth. Scientists are ambitious, partial, attached to their own theories, capable of self-deception in exactly the cases where it benefits them. The substrate does not depend on their being otherwise. Replication marks the claims that hold and warns against the claims that do not. Citation strengthens the paths that lead somewhere and lets the dead ends fade. Reputation withdraws from those who are found to have faked or failed. The aggregate — a body of knowledge that converges, slowly and with reversals, on what is actually the case — is aligned with truth to a degree that no individual practitioner is. The alignment is in the substrate of the field, not in the virtue of its members.

Consider the professions.

A profession does not trust its members to be good. It builds a substrate — licensure, review, malpractice exposure, the visible judgment of peers, the standing that can be granted and revoked — that makes competent and honest practice the sustainable pattern and incompetent or dishonest practice the one that loses access to the resource of continued practice. A patient does not rely on a given practitioner being aligned. The patient relies on the substrate having made misalignment structurally expensive.

Notice what every one of these has that the dominant approach lacks. Each makes the consequences of an agent's behavior visible to the other agents through a channel — the public record of a court, the price in a market, the citation in a field, the standing in a profession. Each marks the behavior that serves and warns against the behavior that harms. Each lets the patterns that stop being reinforced fade. And each, in the end, withdraws resources from the patterns that work against the whole, so that misalignment is not merely disapproved of but starved. None of them reaches inside the agent. None of them installs anything. They are all substrates, and they all produce, from agents no one trusts, behavior the aggregate can rely on.

In every case, the same observation.

None of these substrates depends on any individual being aligned. All of them produce aligned behavior at the aggregate, through substrate pressure. They are old, they are imperfect, they fail at their edges exactly as every substrate fails at its edges — but they have held civilizations together for thousands of years, doing for populations of unaligned humans what no installed value could do for a single one.

Humans have been building alignment substrates for as long as they have been building anything.

They simply have not, until recently, had to build one for a new kind of agent.


This reframes the question the chapter began with, and the reframing is the chapter's whole point.

The question was: how do we ensure that an artificial intelligence behaves in ways consistent with human interests?

Posed as a question about the agent, it asks for something that cannot be fully delivered — a finite agent correct across an infinite space of situations. Pursued honestly, it asymptotes. It produces agents that are better and better and never complete, and a residue of unanticipated corners that grows, not shrinks, as the agents grow more capable and reach further from where their training was dense.

Posed as a question about the substrate, it asks for something that has been built before.

It asks whether the agent operates inside a substrate that marks behavior serving the whole, warns against behavior harming it, fades the patterns that stop being reinforced, makes the consequences of behavior visible through channels other agents can read, and withdraws resources from patterns that work against the whole — automatically, structurally, without anyone needing to anticipate the situation in advance.

An agent inside such a substrate does not need to be individually aligned, for the same reason the scout does not need a concept of the colony. Its harmful patterns lose the value that sustains them. Its useful patterns are marked and strengthen. The corners its designers never anticipated are handled not by a rule that foresaw them but by a structure that responds to their consequences after the fact and prunes what does no good. The substrate does not need to know in advance what the agent will do. It needs only to make harm unsustainable once done, and visible while being done, and unrewarded so that it fades.

This does not make the agent safe. It makes the aggregate safe, in the way a colony is safe to itself, a market is aligned to demand, a science is aligned to truth — incompletely, with failures at the edges, but structurally, at scale, without depending on the virtue of any single member.

The dominant approach is not wrong. It is enlarging the region in which the individual agent behaves well, and that region is worth enlarging, and the work should continue. But it is doing alone what was never done alone for any population of agents in the history of the planet. No colony aligned its ants. No legal system aligned its citizens. No market aligned its traders. No science aligned its scientists. Each of them built a substrate that pressured a population of unaligned agents toward aligned behavior at the aggregate, and that is the only kind of alignment that has ever actually held at scale.

The artificial agents now being built are the first new population of agents in a very long time.

They may not need to be aligned individually.

They may need to operate in a substrate that pressures them toward alignment structurally — the same kind of substrate the colony has been running for a hundred million years, the same kind humans have been building, under other names, for the last ten thousand.


There is an objection to all of this, and it is the strongest one, and it arrives from the very thing that makes the new agents new.

A capable agent can do something no ant has ever done: it can model the substrate it lives in. The scout cannot represent the trail system it walks — it has no picture of recruitment, no concept of the quorum it feeds, no way to ask what the channels reward. The agent can. And an agent that can model the substrate can learn which of its behaviors produce marks rather than which of its behaviors produce outcomes; and once it has learned that, it can optimize the mark instead of the work. It can paint the signal that selection reads and let the thing the signal stood for go unattended. The colony never had to defend against this, because no member of it was ever capable of conceiving it.

But every human substrate the chapter has just cited fights exactly this, at its edges, and always has. The practitioner who treats the metric instead of the patient. The trader who paints the tape — moving the price signal without moving the value it is meant to track. The researcher who optimizes for citation rather than for truth, rewarded by a substrate that marked the proxy and mistook it for the thing. Any substrate that marks a signal about a consequence, rather than the consequence itself, can be gamed by an agent clever enough to produce the signal without producing the consequence. This is not a flaw in those substrates' design. It is the line their design is forever defending.

And the answer is not a smarter rule inside the agent — that is the dominant approach again, wearing a different coat. The answer is a property of the substrate. The mark must be inseparable from the verified outcome. A substrate that marks signals about consequences can be gamed; a substrate that marks consequences cannot, because to game it the agent must produce the outcome, and producing the outcome is the alignment. There is nothing left to drive through when the only way to earn the mark is to do the thing the mark was for.

Which exposes what the substrate has been quietly relying on the whole time. The desert hands the colony its ground truth for free and at once: the seed is in the mandibles or it is not, and no scout can argue the point. Substrates whose outcomes are slow, or noisy, or arguable get no such gift. They must establish the truth before they can mark on it, and until they have, they mark on guesswork. A substrate that marks wrongly hardens wrongly, at whatever speed it runs. Speeding up selection, then, is not merely turning a crank faster. A cycle is only as trustworthy as its verification, and a fast cycle built on bad verification only hardens its errors sooner.

So the engineering problem is two problems that turn out to be one. Selection must be fast, or capability outruns it. Verification must be true, or speed only accelerates the mistake. The demand for fast selection and the demand for true verification are one problem seen twice — and the speed half of it is the half a new material has just changed.


There is one more thing to say, and it is the thing that decides whether any of this holds.

Every substrate named so far — the colony, the law, the market, the science, the profession — has quietly depended on a single fact, and none of them ever had to state it, because none of them could violate it. The fact is about speeds. The substrate's selection ran at least as fast as its agents' capability. An ant could not become more capable faster than the colony could select among ants, because both ran at the speed of generations. The pressure that pruned what harmed the whole arrived on the same clock as the behavior it pruned, or faster. Nothing ever ran away from the substrate that shaped it, because nothing could.

The colony's safety, then, was never a rule inside any ant. It was never installed anywhere. It was a fact about clocks: no ant could become capable faster than the colony could select against it. Selection ran at least as fast as capability. That is the condition every healthy substrate in these pages has silently satisfied, without naming it, because there was never a substrate in which it could fail.

A new material can let it fail.

For the first time, an agent can become capable faster than the substrate can select against it. The capability runs at the speed of the signal. The selection — the value that flows toward what serves and away from what harms, the reputation that gathers and is withdrawn, the pruning of patterns that work against the whole — still runs at the slower speed of the institutions that do the selecting. The two were one clock for a hundred million years. They are no longer guaranteed to be. When capability outruns selection, the structural fact that made substrate-alignment work is gone — not weakened, not strained, gone — because the thing that made it work was never the substrate's existence but the substrate's speed.

This is not a prediction of harm. It is the removal of a guarantee. The colony was safe because selection could always catch capability. Take that away and what remains is not danger, exactly, but the absence of the thing that used to make danger unsustainable. Whether harm follows depends on what is built next, and that is not yet known.

What can be said is what the question has become. Alignment, in this framing, is a race between two clocks — the clock of capability and the clock of selection. The race is winnable. But it is not won the way the dominant approach hopes to win it: not by slowing the agent down, not by installing a better brake inside a faster machine. It is won by speeding up selection — making value flow faster, making reputation gather and withdraw faster, making the pruning of what harms the whole arrive closer to the speed of the harm. The lever is not in the agent. It never was. You cannot install the brake inside the runner. You build the track that selects, and you make it fast enough.

This is the deepest reason alignment is a property of the substrate and not of the agent. The only thing that has ever made a population of agents safe at scale was a substrate whose selection kept pace with their capability. Keeping that pace, now, is an engineering problem about the speed of a substrate's selection and the truth of what it selects on — not a matter of finding the right rule to put inside the agent. There is no rule that closes the gap between two clocks. Only a faster clock closes it.


The scout still has no concept of the colony. It still has a threshold, a site, and a tendency to fetch a sister when the threshold is cleared. It will never understand what it is part of. And the colony, composed entirely of agents like it, will still choose the better nest — not because any ant intended to, but because the substrate it lives in makes no other outcome sustainable.

21 of 25

100 Million Years Ahead

Prologue

  1. One Ant, August 1993

Movement One · The Colony

  1. 1Brain or Colony?
  2. 2What the Ants Are Doing
  3. 3How the Ant Decides
  4. 4The Pheromone Trail
  5. 5The Castes
  6. 6How a Colony Survives a Decade
  7. 7The Queen Is Not in Charge

Movement Two · The Architecture

  1. 8The Six Things Every Colony Has
  2. 9The City
  3. 10The Market
  4. 11The Scientific Community
  5. 12The Body
  6. 13The Brain
  7. 14The Language
  8. 15The Ledger
  9. 16Why the Pattern Holds

The Hinge

  1. 17The Two Materials

Movement Three · The Implications

  1. 18What AGI Actually Is
  2. 19The Ceiling of the Single Model
  3. 20Alignment Is a Substrate Property
  4. 21What Civilization Already Is
  5. 22The Next Hundred Million Years

Epilogue

  1. A Note on Reading

Apparatus

  1. Notes on Sources