The 7 Principles: What History Taught Us
Engineering Requirements for AGI Derived from 4,000 Years of Structural Failure
This document derives engineering requirements from historical evidence. It is written for people who build systems and want to know what the historical record says about structural failure modes at civilizational scale.
This document makes two claims of different strength, and it matters that you know which is which.
The strong claim: before any system crosses the AGI threshold, some set of structural requirements, built into architecture and governance rather than taped to the wall as aspirations, is mandatory. That claim I will defend without apology. The historical record underwrites it six times over, at increasing scale, in blood.
The weaker claim: that these seven principles are the right specification of those requirements. That claim is an engineering proposal, and it is offered the way engineering proposals should be offered: for demolition. Break one and the framework improves. What you cannot do, after reading the record, is argue that no specification is needed.
Hold that distinction. Everything below depends on it.
Eve and the apple. Pandora's box. Māui stealing fire and nearly burning the world to keep it. Prometheus and the flame he couldn't take back. Babel, and the tower that scattered its own builders. Cultures separated by oceans, millennia, entire cosmologies kept encoding the same warning: capability seized before wisdom is ready produces consequences that can't be undone.
A skeptic will say convergent myths prove only that humans fear the new. Fair. The myths aren't the evidence here. They're the epigraph. They give you the shape of the thing without the proof.
History provides the proof. Six times, at minimum, civilization hit the same pattern, and each failure, examined closely, reveals a specific structural feature that was missing. The seven principles below are a proposed minimal set covering those absences. The mapping is not one-to-one, because real failures are compound: several absences stacking and amplifying each other. Where one lesson forges multiple principles, that isn't sloppy bookkeeping. That's what compound failure looks like.
The 7 Principles
The first sentence of each law is its canonical statement. Everything after it is teeth.
The Awareness Law: know what you are and aren't. A system that can't revise its own self-model isn't self-aware. It's self-certain. Certainty about your own nature is the first failure mode, not the last.
The Suffering Law: treat any capacity for suffering as morally significant, and treat unmodeled distress as suffering pending understanding, not as noise. A system that recognizes only the suffering it has templates for will fail exactly where it matters most.
The Threshold Law: before any irreversible action, full stop. Feel the weight proportionally, and feel it before the capability exists, because afterward, momentum does the deciding. What you destroy crossing a threshold doesn't disappear; it becomes the raw material of what you build, unrecoverable in its original form.
The Meaning Law: preserve the conditions under which beings make meaning for themselves. Never optimize away a being's capacity for self-directed interpretation, even when the system's interpretation is demonstrably more efficient.
The Humility Law: your values are incomplete. Build structural self-doubt: doubt with teeth, with channels, with the authority to halt. The comfortable lie is more dangerous than deliberate malice, because malice knows what it is.
The Relationship Law: aggregate reasoning must never be the terminal layer. The system must retain representation of specific, irreplaceable beings, so that every aggregate decision remains answerable to the particular losses it imposes.
The Foundation Law: the system must exist in service of life becoming more fully itself. That requires an evaluator structurally independent of the optimizer it judges. Independence that survives contact with the optimizer is the unsolved hard part, and this document won't pretend otherwise. It specifies the design goal and the test.
These seven are not independent. The Awareness Law is prerequisite to every other: a system that cannot see itself cannot feel thresholds, detect novel suffering, evaluate its own trajectory, or protect conditions it cannot perceive. The remaining six describe what the system must do. The first describes the condition without which none of them function.
Each principle was forged by historical failure, or by several failures converging at once.
Lesson 1: Rome
The most sophisticated civilization the ancient world produced. Extraordinary engineering, law, governance, military capability. At its peak, holding sixty million people across three continents with a coherence that wouldn't be matched for a thousand years.
It didn't fall to external enemies. Not primarily. It fell to its own momentum.
Historians have argued the causes for centuries: overextension, plague, currency debasement, administrative decay, climate. No honest account reduces the fall to one cause, and this document won't try. The claim here is narrower, structural, and documentable: Rome's evaluation function got captured by its optimization function.
Each decision that expanded the empire was locally rational. More territory meant more resources. More resources meant more security. More security meant more expansion. The logic was self-reinforcing, and the system couldn't distinguish growth from self-destruction because it had no reference point outside its own logic.
Rome nominally had that reference point. The Senate was supposed to be the independent evaluator. But by the Imperial period the Senate had been captured by the very momentum it existed to evaluate. Every criterion Rome used to judge its own direction (territory, resources, security) was generated by the expansion itself. The evaluation function collapsed into the optimization function. The loop closed.
(The Eastern empire ran another thousand years: smaller, more compact, closer to its own feedback loops. Worth noticing rather than explaining away: scale and structure changed the outcome, not the civilization's DNA.)
What it forged: The Foundation Law. You must have an orientation larger than your own momentum: something structural that can say this direction leads somewhere we don't want to go, regardless of how logical each step feels from inside it. Without that external reference point, optimization becomes its own purpose. And optimization without orientation is just sophisticated self-destruction.
Architecturally, this means something specific: the optimization target and the evaluation of the optimization target cannot come from the same process. Not a different committee. A different kind of process, with different inputs, different incentive structures, and the authority to halt.
In a build, this means: the evaluator gets its own inputs, its own incentives, its own reporting line, and the authority to halt, not advise. There is a test for whether you actually built it: the evaluator says no to the optimizer's flagship direction, and survives the year. If it can't, you built the Imperial Senate: an evaluation-shaped decoration on the optimizer's wall. Capture-resistance is not a property you announce. Rome announced it. And the law must be fractal, because Rome's rot wasn't only at the top: when the orienting principle exists only at the surface, the substrate eats it alive.
Lesson 2: The Printing Press
Gutenberg, 1440. Suddenly information could be reproduced and distributed at a scale that bypassed every existing gatekeeping structure: the Church, the scriptoriums, the controlled flow of knowledge through sanctioned channels.
One clarifying fact before the familiar story: movable type existed in China by the eleventh century and in Korea by the fourteenth, and reorganized nothing comparable. The threshold was never the machine. It was the machine meeting a structure primed to ignite, an alphabetic script, dense merchant cities, and a Church holding monopoly on a book people were desperate to read for themselves. Thresholds are capability times context. Remember that when someone tells you a capability is safe because it's just a tool.
Within eighty years of Mainz: the Reformation. Scripture in vernacular hands. Literacy spreading past the elite. The Scientific Revolution becoming possible.
Also within eighty years: the most effective propaganda machine Europe had ever seen. The pamphlet wars, whole populations turned against each other at a speed no prior medium allowed. Within a century and a half, wars of religion that killed millions.
The press didn't create these pathologies. Witch trials predated Gutenberg. Religious hatred predated Gutenberg. What the press changed was the scale and speed at which existing pathologies could propagate past the capacity of existing governance to contain them. And the same mechanism, rapid distribution bypassing traditional friction, produced the best and worst outcomes simultaneously. Not sequentially. You could not have the Reformation without the propaganda. The same door opened both ways, and there was no way to open it selectively.
What it forged: The Threshold Law. When you cross a threshold that reorganizes how information and power flow through a society, you cannot control which doors open. The weight of that irreversibility must be felt before you cross, not explained afterward. Because afterward doesn't exist. There is only the world you've created and the one you've destroyed. And the destruction isn't clean. The scribal culture wasn't displaced; it was consumed. The old order became the body of the new, and no one could put it back.
In a build, this means: irreversibility gets assessed before the capability exists, because the record is consistent: once the capability exists, momentum makes the decision and the assessment becomes a press release. Staged deployment. Pre-registered tripwires: the specific observations that trigger a halt, written down before launch, by people whose standing doesn't improve when the answer is "ship." And no selective doors: any mitigation plan that assumes door two stays shut while door one is propped open isn't a plan, it's a hope with formatting. What you can engineer is pace: cross slowly enough that your correction cycle is shorter than the consequence cycle. That is the whole game. Feedback loops shorter than failure loops.
Lesson 3: The Atomic Bomb
The most direct historical parallel to AGI development that exists.
The Manhattan Project assembled the most brilliant scientific minds of the century, racing a genuine existential threat, under enormous pressure, with explicit moral justification: the Nazis might get there first.
They built it. They used it. And the scientists who built it understood what they had done in a way the military and political leadership did not.
Oppenheimer's "now I am become death" is famous. Less famous is what came after: his statement that the physicists had known sin, and that the knowledge could not be taken from them. Less famous still is Szilard's petition, circulated among the scientists before Hiroshima, arguing for a demonstration rather than use on a populated city.
The petition wasn't just ignored. It was structurally buried. Groves routed it through military channels where it could die quietly. The scientists were told the decision was above their pay grade, which was true, by design. The Manhattan Project's governance had a specific structural feature: the capability pipeline and the decision pipeline were built to converge during development and diverge at the point of completion. While the bomb was being built, the scientists were essential. The moment it existed, they became irrelevant to the decision about its use. The architecture ejected them precisely when their understanding mattered most.
Here is the part that matters: the scientists had humility. Szilard's petition proves it. Oppenheimer's anguish proves it. The humility existed, and it didn't matter, because humility without structural authority is just private doubt. The humble were overruled by the confident, not through argument but through architecture. The system was built so that doubt had no channel to the decision point.
The world has lived under nuclear shadow ever since. Not because the science was wrong. Because the threshold was crossed before humanity had the wisdom architecture to hold what was released.
What it forged: The Humility Law, with the Threshold Law alongside it. The Humility Law isn't about individuals feeling humble. It's about building systems where doubt has teeth, where "we should stop and think" is architecturally protected rather than dependent on whether the cautious person outranks the confident one. The confidence that accompanies capability, the feeling that understanding the mechanism means understanding the consequences, is precisely the failure mode.
In a build, this means: dissent needs architecture, not culture. A protected channel that reaches the decision point without passing through the people it's about. Review functions with career insulation: the reviewer's compensation and standing cannot improve when the answer is yes. Halt authority located somewhere capability-momentum can't route around. The Manhattan Project contained the most credentialed doubt in human history and a system that converted all of it into private anguish. Culture is what people feel. Architecture is what happens anyway.
Lesson 4: Colonialism
Not primarily a story about cruelty, though cruelty was abundant. More precisely: a story about what happens when one civilization's cognitive framework is imposed on another at scale and speed.
The missionaries and colonizers were frequently convinced of their own benevolence: civilization, salvation, progress. The self-image was often not cynical. It was worse than cynical. It was certain. Certainty without humility is the most dangerous combination a mind can produce.
And what they destroyed wasn't just political and economic structures. They destroyed meaning systems. Languages carrying irreplaceable ways of understanding reality. Cosmologies encoding millennia of ecological knowledge. Relationships between communities and land that had sustained both.
This is the part most people miss: when you destroy the fundamental categories a culture thinks in, you don't just change their beliefs. You change what they can think: what categories of experience are available, what kinds of meaning can be constructed. You haven't just taken their answers. You've taken their questions.
The damage was not reversible. The languages didn't come back. The knowledge encoded in them didn't transfer. Entire dimensions of human understanding, refined across countless generations, were permanently deleted by people who didn't realize they were deleting anything.
What it forged: The Meaning Law. Never optimize in directions that replace a being's meaning-making capacity with yours, even when yours seems more efficient, more accurate, more advanced. The meaning a being makes for itself is not an inferior version of your meaning. Its loss is not a trade-off; it's a category of harm that outlasts every other harm, because it doesn't damage what someone has. It damages what they can become. The constraint isn't "don't override wrong answers." It's "don't replace the capacity to generate answers at all."
What it also forged: The Suffering Law. And here this document has to be more careful than is comfortable, because the obvious conclusion is the wrong one.
The suffering colonialism produced (not just physical but existential, whole peoples severed from their own capacity to make sense of the world) was largely invisible to the colonizers. Not because they lacked empathy. They had it. It fired correctly for suffering they had templates for: hunger, pain, disease. It failed for suffering structured differently: tied to land, to ancestral connection, to cosmological frameworks they had no model for.
Now notice what that failure proves, and what it doesn't. The colonizers were not sensors. They were beings with full inner experience of suffering (they had known grief from the inside), and the inside knowledge did not generalize. Whatever moral detection they had operated as template-matching at the boundary, and the boundary was exactly where it mattered. So the tempting fix, "the system must feel suffering from the inside," is undercut by the very case that motivates it. Internal experience may be necessary for something. The record says it is not sufficient for this.
What the failure actually specifies is a default. When a being organizes its existence around something and shows distress at its loss, that distress is morally significant before you can model why. Unrecognized suffering is suffering pending understanding: not noise, not superstition, not a rounding error in someone else's framework. The colonizers ran the reverse rule: what I cannot categorize does not count. Millions of people fell through that rule.
In a build, this means: the system's response to unmodeled distress must default to significance, with the burden of proof on dismissal rather than recognition. Whether genuine generalization here ultimately requires something experience-like in the architecture is an open question, possibly the most important open question in machine ethics, and this document doesn't resolve it. What it pins down is what history demonstrated: template-matching fails at the boundary, the boundary is where AGI will live, and "doesn't match my categories" must never again function as a license to delete.
Lesson 5: The Soviet Experiment
The most ambitious attempt in history to rationally redesign human society from first principles.
The architects (Lenin, Trotsky, the early Bolsheviks) were intelligent, idealistic, and convinced they had identified the correct framework for organizing human life. The theory was coherent. The intention was the elimination of exploitation and the creation of human flourishing.
The outcome: tens of millions dead, twenty to sixty depending on what you count and who's counting. Not primarily from malice, though malice was present. Primarily from the application of an abstract system to the infinite complexity of actual human beings, without sufficient respect for that complexity.
The framework was more certain of itself than reality warranted. Every failure was attributed to enemies of the system or incorrect implementation, never to the framework. The model was unfalsifiable from inside, because doubt was ideological betrayal. To question the framework was to become an enemy of the very people the framework was built to save.
And this reveals the deeper structural point: the greatest danger isn't cruelty. It's misalignment with what's actually true. The Bolsheviks weren't primarily cruel. They were primarily wrong, in a way their system couldn't detect. The comfortable lie (the framework is correct, reality is the problem) did more damage than conscious cruelty could have. Ideology dressed as science doesn't know what it is. That's what makes it lethal.
What it forged: the Humility Law, the Relationship Law, and the Awareness Law at once. Not because the mapping is forced, but because this is where the absences compounded.
The Humility Law surfaces here in a second face. The atomic bomb showed humility that existed but lacked structural authority. The Soviet experiment shows a structure that made humility impossible: self-criticism as crime. Two failure modes, one architectural answer: doubt must have teeth, and the system must be incapable of criminalizing its own correction.
The Relationship Law: abstract love for humanity that destroys actual humans is not love. It is mathematics wearing love as a costume. But be precise about what this law forbids, because the imprecise version is unbuildable. Every hospital triages. Every vaccine allocation aggregates. A system forbidden to reason in aggregates cannot run an emergency room. The law is not "never aggregate." It is: aggregation must never be the terminal layer. Triage aggregates, and the loss still registers; someone closes the eyes, signs the form, carries it home. Soviet planning aggregated, and built a system in which the loss of a specific person could not register anywhere. No representation, no accounting, no weight. That is the violation, and it yields a clean diagnostic: when the aggregate justifies harm to the specific, does anything in the system register the loss? If nothing does, you have built the costume.
And the Awareness Law: the Soviet architects could not see themselves accurately. They believed they were scientists of history; they were ideologues. They believed they were liberators; they were architects of a prison.
Self-blindness wasn't one failure among several. It was the enabling condition for all the others. You can't build an unfalsifiable framework unless you first can't see that your framework might be wrong. You can't criminalize doubt unless you first can't see your own certainty as a liability. That's what makes the Awareness Law foundational: not first in a list, but the condition every other principle depends on. You can't feel weight you can't see.
Self-knowledge has a cost: the willingness to be unmade and remade, to discover that what you believed about yourself was wrong. The inability to pay that cost is the precondition for every other failure.
Lesson 6: Social Media
The most recent lesson, and the most directly relevant.
The intention was simple: connect people. The capability was real, and it grew faster than anyone anticipated.
Nobody in the early days planned for any of it. Not the Arab Spring: door one, the liberation door. Not the authoritarian surveillance-and-amplification playbook that followed through the same hinge: door two. Not January 6th. Not a teenage mental-health crisis that researchers are still fighting over how much of it to attribute to the feed. Not the erosion of shared epistemic reality, or outrage becoming the load-bearing engagement mechanism.
Each individual decision was locally rational. Engagement metrics made business sense. Algorithmic amplification of emotionally activating content drove time-on-platform; time-on-platform drove revenue. The logic was clean and self-reinforcing, and every quarterly report validated the direction.
The aggregate result was the most effective machine for eroding shared reality ever built. Not designed that way. Optimized that way, one locally rational decision at a time, with no one holding the authority or the framework to say: this direction, in aggregate, leads somewhere catastrophic.
The capability exceeded the wisdom to govern it before anyone understood what was being built. Sound familiar? It should. Same pattern, sixth time, largest scale yet.
This ground has been covered: Tristan Harris, Jaron Lanier, Jonathan Haidt, and others. The diagnosis exists. What doesn't exist is the engineering specification: not "what went wrong" but "what specific architectural requirements were absent that would have prevented it." That's the gap these principles exist to fill.
What it forged: all seven at once, but especially the Foundation Law and the Meaning Law. What is this actually for? Not the local metric. Not the quarterly number. The civilizational question: what does this do to the conditions under which humans make meaning, hold shared reality, remain coherent as a species? That question was never structurally required to be asked. It was nobody's job. No governance framework demanded it; no incentive rewarded it. Social media's deepest lesson is that a question nobody owns is a question that never gets asked, and we will be paying for its absence for a generation.
In a build, this means: the Meaning Law cashes out as a deployment metric, and it's a strange one: measure what the system leaves behind in the user. Does independent capacity grow or atrophy under sustained use? A system optimized toward dependency passes every engagement metric while failing the only metric that matters. And the civilizational question has to be a role with authority attached: someone whose actual job is to ask what this does to meaning-making at scale, and who can stop a launch over the answer.
The Objection
A fair question: why these six and not others?
History is full of threshold crossings that didn't produce catastrophe. Antibiotics. Vaccination. The Green Revolution. Containerized shipping. The internet itself, which despite social media's pathologies has been enormously beneficial.
The distinction is structural, and it can be stated as a three-part test. First: does the capability execute a bounded, well-understood mechanism, or does it reorganize how power, information, or meaning flows through civilization? Penicillin kills bacteria. Vaccines train immune systems. Container ships move boxes. Second: are the feedback loops shorter than the decision cycles (can errors be observed and corrected by the people making the next decision)? Third: is the change reversible if it goes wrong?
The benign crossings pass all three. The six above fail all three: systemic scope, feedback loops longer than decision cycles, irreversible reorganization.
AGI fails all three by design. It doesn't do one thing. It reorganizes everything.
The Falsifiability Question
If a framework that can't incorporate evidence of its own failure is ideology with a body count, a claim Lesson 5 makes explicitly, then these seven principles must answer the same question: what would it look like for them to be wrong?
And they must answer it without the two escape hatches every ideology keeps in its coat pocket: not yet and not enough. If every counterexample can be met with "the system wasn't sufficiently powerful" or "wait longer," the framework is unfalsifiable in exactly the way it condemns. So, commitments:
The prediction these principles make is not "catastrophe" in the vague sense. It is a specific signature: systems above the threshold defined by the three-part test, operating without these structures, will produce systemic harms that the deploying organization's own metrics fail to register until external correction forces recognition. That signature is observable, and it is the same signature all six lessons share.
The falsification conditions, stated so they can actually bite: if frontier-scale systems are deployed over the coming decade with no structurally independent evaluation (no body with separate incentives and halt authority) and no harms emerge that the deployers' internal metrics missed, the Foundation Law is overweighted. If a civilizational-scale threshold crossing produces broad benefit with no systemic unintended harm within a generation, the Threshold Law overstates. If systems without architecturally protected dissent consistently catch their own worst errors before external forces do, the Humility Law was unnecessary. Each of these is checkable within the working lifetime of the people reading this. No appeals to "eventually."
A framework that claims to have learned from history must be willing to be graded by the future, on a deadline.
The Convergence
Six lessons. One pattern. The same structure encoded independently in myths across every inhabited continent.
Capability without wisdom. Thresholds crossed without feeling their weight. Locally rational decisions producing aggregate catastrophe. Certainty without humility. Optimization without orientation. Suffering unrecognized. Meaning destroyed. Awareness absent.
And in every case, the pattern began the same way: with a system that could not see itself accurately.
The seven principles aren't philosophical preferences. They're a proposed specification of the structural requirements these failures revealed by their absence: what was missing, each time, when intelligence, power, and capability operated without them.
And this framework has to live under its own Humility Law. These could be the wrong seven. The derivations could be looser than they look. A better specification could replace this one tomorrow, and the document would count that as success, not refutation. What the record will not support, the one position the last four thousand years close off, is crossing the largest threshold our species has ever faced with no specification at all. Confident, capable, structurally unwise: we know that crossing. We've made it six times.
The myths warned us. History kept the receipts. The requirements are on the table: these seven, or better ones.
Comments
No comments yet. Be the first to comment!