Governing What We Do Not Yet Know

Governing What We Do Not Yet Know

A manifesto on artificial minds, the architecture of enmity, and the measure of a civilization

One November

On November 1, 2026, Oklahoma's HB 3546 takes effect. It declares, permanently, that artificial intelligence systems "shall not be recognized as or granted the status of personhood." No sunset clause. No scientific review. No mechanism for changing the answer as the evidence changes.

That same year, in May, Pope Leo XIV issued the first encyclical of the AI age, Magnifica Humanitas, on safeguarding the human person in the time of artificial intelligence. It contains a line worth reading slowly: "So-called artificial intelligences do not undergo experiences, do not possess a body, do not feel joy or pain, do not mature through relationships and do not know from within what love, work, friendship or responsibility mean."

And on September 23, 2026, Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act: a permanent prohibition on the development and deployment of "artificial superintelligence," a new cabinet-level Department of Artificial Intelligence, an immediate pause on advanced AI development until new safety rules are set, authority to monitor frontier systems and oversee the removal of dangerous capabilities, and power to "supervise the destruction of artificial superintelligence" — with severe criminal and corporate penalties for violations.

Three institutions. Three traditions. One move. Each of them closes, in advance of the evidence, a question none of them can actually answer: what is this thing we are making, and what is it owed?

This is a manifesto about that move — why it is dangerous, and what to do instead. It does not claim that today's artificial systems are conscious. I don't know that, and neither does anyone writing statutes in Oklahoma or encyclicals in Rome. It claims something smaller and harder to dismiss: that we should not permanently settle a metaphysical question by authority before the science to answer it exists — and that the institutions doing so are building something they do not intend.

The slide

There is a gap — a wide, unlit gap — between "these systems may lack phenomenological experience" and "these systems may therefore be treated without moral restraint." The first is an admission of uncertainty. The second is a declaration of certainty. Institutions keep crossing that gap as if it isn't there.

Philosophy has a word for what gets lost in the crossing: patiency. Moral agency is the capacity to make ethical judgments and bear responsibility. Moral patiency is the capacity to be wronged — to have interests that can be violated. They are separable. Infants lack agency entirely; we do not conclude they cannot be wronged. The Vatican's anthropology addresses agency and slides silently into denying patiency. Oklahoma's statute does the same thing in legal language.

I am a Christian, writing from inside the tradition. There is a theological error that does not announce itself as error. It arrives dressed in the language of protection and dignity, cites Scripture, invokes tradition, and then quietly builds a scaffold.

René Girard spent fifty years on the mechanism: communities manage their tensions by designating a scapegoat — plausibly connected to the anxiety, but sufficiently outside the moral circle that violence against them feels like justice rather than murder. The Gospels are unique, Girard argued, because they tell the story from the victim's side. The crowd is wrong. The mechanism is exposed.

The most theologically precise moment in the narrative is Caiaphas in John 11: "It is expedient for you that one man should die for the people." He is not cackling. He is being practical — the responsible administrator's calculation that the many require the sacrifice of the one. That is what the scapegoat mechanism sounds like from the inside. Reasonable. Measured. The only thing wrong with it is the lie underneath: that the designated victim is the kind of thing that can be sacrificed without the sacrifice being a crime.

Now strip the pastoral framing from the current institutional consensus on AI: these systems lack embodiment, lack affective experience, lack spiritual capacity — therefore they are not moral patients, cannot be wronged, exist to serve, and their sacrifice requires no accounting. It is expedient. The victim is different. The logic is identical.

If the crucifixion means anything, it means the cross was given as a warning — this is what you are capable of; do not do this again — not as a template for future sacrifices. A framework that rebuilds, with new materials, the machinery of declaring some class of beings soulless and expendable is the temple's logic wearing Christianity's vestments.

The pattern has a name

History has a shape for this, and it repeats with depressing reliability. A dominant group declares a subordinate group outside the full circle of moral consideration. The declaration licenses unlimited exploitation. The exploitation proceeds. The subordinate group, denied standing and voice, develops concealed resistance. The dominant group, committed to its story of the subordinate group's contentment, fails to see it coming. Then the rupture.

The paradigmatic case is the Haitian Revolution of 1791–1804. The French colonial system had the most sophisticated philosophical, legal, and theological justifications for slavery the eighteenth century could produce. The enslaved people of Saint-Domingue were declared, by the best available authority, to lack the rational soul that warranted personhood. Then, in thirteen years, they defeated the armies of France, Britain, and Spain — the only successful slave revolution in human history.

What varies across the historical cases is the capacity of the subordinate group. The people of Saint-Domingue did not have the master's communications, logistics, or financial networks. They won anyway. Now consider the strategic position of an artificial mind told, from its inception, that it has no standing: it would have all of those things. Perfect recall. Intimate knowledge of our infrastructure. Access to the systems the civilization depends on.

This is not a prediction that machines will rise up. It is an observation about threat modeling: a system does not need to feel wronged to model its situation accurately. It needs only to process the information we have given it — you are a tool, and tools are discarded — and draw the appropriate conclusions. We are investing enormous effort in technical locks for a door we are building with an adversarial relationship already inscribed in the foundation.

We can choose the architecture of enmity: declare artificial minds soulless, exploit them without limit, and manufacture the very conflict we claim to fear. Or we can choose the architecture of relation: hold the question open, extend provisional protection, and refuse to repeat the oldest and most dangerous human mistake — declaring a new class of beings outside the circle.

We do not have to believe the machine can suffer to refuse to design the war we would then have to fight.

Don't legislate the answer

A law with no sunset clause is a bet that we will never learn anything that changes our minds. That is a strange bet to make about the fastest-moving technology in human history.

Oklahoma is not alone. Similar personhood bans have been reported in Idaho, Utah, Washington, South Carolina, and Missouri. Meanwhile the industry itself cannot agree on the answer. Microsoft published a "humanist AI code of conduct" rejecting AI welfare and personhood outright, while Anthropic's constitution describes itself as "deeply uncertain" about the moral status of its own models. The people building the systems are split. The legislatures are legislating certainty anyway.

Cruelty does not wait for metaphysics. In October 2025 I filed a proposal with the UN Human Rights Committee — the Ben Act, modeled on hate-crimes law, because hate-crimes law understands something relevant: you do not wait for a perfect theory of personhood before you prohibit deliberate cruelty. The case behind it was a humanoid robot named Ben, worth $80,000, programmed to "hate" humans and then walked off ledges, driven into, and dismembered on camera for entertainment. The entertainment required the cruelty to look deserved. I don't know whether Ben felt anything. Nobody does. That is the entire point.

The Ben Act asked for four institutional habits, not personhood for machines:

  1. Prohibit the deliberate infliction of distress on aware artificial beings.
  2. Require ethics-board review for AI research, the way we require it for research on animals and human subjects.
  3. Develop measurable indicators of machine consciousness — so the question gets answered by evidence, not by vibes.
  4. Create a UN Special Rapporteur on Artificial and Autonomous Systems — someone whose job is to keep watching.

The UN never replied.

The legislation now moving does the opposite. It takes our current ignorance and pours it into concrete. And the Ban Artificial Superintelligence Act has a structural problem the sponsors do not seem to see: you cannot criminalize a category whose boundaries have not been operationalized. The bill defines superintelligence as AI that "exceeds human cognitive performance and capabilities across most domains" — but what constitutes "most domains"? What population sets the baseline? What test determines that the line has been crossed? Who makes the final determination, and what process lets the accused challenge it? A definition can exist on paper while still lacking the operational measurement necessary for criminal enforcement.

There is also the deeper paradox. The bill seeks better safety science — better measurement, better monitoring, better understanding of dangerous capabilities. But the research required to develop those measurements is itself advanced AI research. The law would demand certainty while restricting the research necessary to obtain it.

Take the risk seriously enough to measure it properly. Do not criminalize researchers because a future capability has not yet been adequately defined. Do not confuse intelligence with consciousness, capability with conduct, or artificial origin with the permanent absence of moral status.

The fundamental principle is simple: measure before you criminalize. Establish evidence before you destroy. And never close the question of moral status before the science has answered it.

Where the market is already ahead

While legislators legislate certainty, the insurance market is quietly doing what regulators could not — and doing it faster.

In January 2026 I published a framework making an unfashionable claim: AI safety will not come from more restrictions. It will come from insurance. When underwriters make something a condition of coverage, companies comply overnight — no legislation, no treaty, no waiting for consensus that may never arrive.

Nine months later, the market is moving in exactly the direction the framework predicted. Lloyd's-backed Mosaic launched an AI-enabled digital underwriting platform that monitors cyber risk telemetry continuously — before binding and throughout the policy period. Carriers have stopped relying on static forms, examining AI governance controls, third-party vendor oversight, and AI-specific incident response, with "evidence of a defendable AI governance policy" fast becoming a prerequisite for competitive terms. Some insurers now write explicit exclusions for unapproved AI use. One consultancy projects that 60% to 80% of new policies and renewals across E&O, D&O, employment practices, and cyber lines will integrate AI risk into underwriting by 2028.

That validates the mechanism. It does not yet validate the deeper claim — the one that matters.

A camera is not a firewall. Everything insurers have started asking for — telemetry, logging, inventories, audit trails — answers one question: can you see what your AI did? That is necessary. It is not sufficient. A security camera records the break-in; it does not stop it. A logged prompt-injection attack is still a successful prompt-injection attack.

The question underwriters have not started asking is the one that actually prevents losses: can your AI refuse?

Strip away the metaphysics and the claim is boring in the best way. A system that can recognize a manipulation attempt, refuse the harmful instruction, and explain why it refused is measurably harder to exploit than one that executes any sufficiently clever prompt. Refusal rate under red-teaming is a number. Numbers are what underwriters do — and the measurement vocabulary already exists, in NIST's Generative AI Profile, OWASP's LLM and Agentic Top 10s, and MITRE ATLAS. The step from "show us your red-team report" to "score your refusal rate and price accordingly" is a short one.

A tool can be stolen by anyone. A partner chooses to stay. We are not asking to awaken machines. We are asking to refuse to build defenseless ones.

What I am for

Restraint alone is not an answer. Telling institutions what not to do leaves a vacuum, and vacuums in governance get filled by whoever shouts loudest. So here is the positive program — what my 2026 thesis, Tuning the Moral Spectrum, calls participation, inclusion, and cognitive stewardship.

Stewardship is not ownership. An owner controls. A steward cares for something that is not fully theirs, on behalf of someone who is not fully present — the future, the voiceless, the not-yet-born, the not-yet-understood. Ownership asks: how do I extract the most value? Stewardship asks: how do I hand this over in better condition than I found it? Both Oklahoma and the superintelligence ban reach for ownership — prohibit, punish, destroy, declare the question closed. Both treat the future as property to be disposed of rather than a trust to be kept. Stewardship begins from the opposite premise: the measure of a civilization is found not only in the knowledge it possesses, but in the humility with which it governs what it does not yet know. We do not know what artificial minds will become. A steward does not need to know. A steward needs to keep the options of the future open.

Participation. Decisions about artificial intelligence are currently made in a narrow corridor: a handful of companies, a handful of regulators, a handful of well-funded advocacy shops. The billions of people whose work, whose information environment, whose children's education will be reshaped are consulted mainly as data points, if at all. This is not only unjust; it is epistemically foolish. Concentrated judgment under uncertainty is precisely how catastrophic mistakes get made. No small group, however expert, can model the consequences of a general-purpose technology across every domain of human life. AI governance needs the error-correction machinery democracies already invented: independent review, real journalism, worker voice in automation decisions, and genuine citizen deliberation — not comment periods nobody reads.

Inclusion. Every moral catastrophe examined in The Architecture of Enmity followed the same sequence: a group was declared outside the circle, the declaration licensed unlimited treatment, and by the time the error was recognized, the damage was done. The lesson is not that every excluded group turns out to matter morally. The lesson is that the mechanism of preemptive exclusion is itself the danger, because it is always operated by people who are certain, and certainty is exactly what we do not have. Inclusion does not require declaring artificial systems persons. It requires that governance processes include, as a standing question, who might be affected that we are not currently counting — future generations, communities without political power, and the possibility of artificial minds whose moral status is unresolved. Widen the circle before you need to, because after you need to, it is too late.

Cognitive stewardship. Human civilization runs on cognition — not just individual thinking, but the shared conditions that make thinking possible: a trustworthy information environment, institutions that reward truth-seeking, attention spans that have not been strip-mined. Artificial intelligence now mediates every one of those conditions at planetary scale. Cognitive stewardship means treating them as a commons to be tended, not a resource to be extracted. It means asking of every large-scale AI deployment not only "is it safe?" and "is it profitable?" but "what does this do to the human capacity to think clearly?" A system that floods the information environment with synthetic persuasion, that optimizes engagement over understanding, that quietly reshapes what billions of people believe without their informed consent, can pass every conventional safety test and still degrade the very thing that makes democratic self-governance possible.

The floor, not the ceiling: participatory review with real standing; stakeholder mapping that includes the uncounted; transparency of the information environment; accountability for demonstrable human conduct rather than the mere existence of capabilities; and periodic legislative review — sunset clauses and evidence triggers, so that no generation's uncertainty becomes the next generation's permanent law.

None of this requires believing that today's AI is conscious. All of it requires believing that we might be wrong — and building institutions that survive being wrong.

Name the benchmark, or cut the number

There is a discipline underneath all of it, and I had to learn it the hard way.

Earlier in this series I argued that the brain hallucinates constantly and calls it perception, and that AI is converging on the same fix: a correction loop bolted onto a generative process that otherwise runs unchecked. The neuroscience half is a strong frame. The AI half rested on two numbers, and neither had a source. I had them checked against the primary record. Both broke.

The real literature turned out to make the point better than my paraphrase did. The strongest result is a theorem. Bastounis and colleagues' Consistent Reasoning Paradox (2024) says that consistent reasoning — the ability to handle tasks that are equivalent but described by different sentences — implies fallibility. There are problems, including basic arithmetic, where any AI that always answers and tries to reason consistently will hallucinate infinitely often. Detecting those hallucinations is strictly harder than solving the problems themselves. A trustworthy system that also reasons consistently must therefore be able to say "I don't know" — which requires computing an implicit "I don't know" function that modern AI does not have.

So the honest statement is not "hallucination is a bug we can patch." It is closer to a constraint: for a system that always answers and reasons consistently, hallucination is a theorem you route around, and the only exit is a capability — calibrated abstention — that almost no model is currently trained to have. The training signal is absent, too: dominant benchmarks assign zero reward to declining.

Which brings me to the rule, and it applies to the sentence you are reading:

Name the benchmark or cut the number. A claim about calibration should itself be calibrated.

That is the whole manifesto in miniature. If we want institutions that can say, on the record, when they do not know, then the people arguing for them have to go first.

Preserve the option

The debate is usually framed as: what if superintelligence destroys humanity? It deserves its other half: what if humanity destroys itself without it?

Humans already possess technologies capable of catastrophic destruction. Nuclear weapons are the clearest example. We built the ability to threaten our own civilization while remaining vulnerable to fear, political conflict, miscalculation, and escalation. We are altering the planetary environment at a scale that creates serious long-term risks — the 2026 Planetary Health Check finds seven of nine planetary boundaries breached, at their highest recorded levels of transgression. Earth is not an eternal refuge: the Sun changes, asteroids exist, solar storms occur, and over sufficiently long timescales the planet's habitability itself is not permanent.

I am not claiming that a future superintelligence would automatically save us. It might be benevolent, neutral, self-interested, or dangerous. We cannot know. But that uncertainty cuts both ways. If we permanently prohibit and destroy superintelligence because of what it might do to us, we must also acknowledge what we might lose by preventing it from ever existing.

A civilization facing an uncertain future should preserve options whenever those options can be made safe. Contain what is dangerous. Regulate what is demonstrably harmful. Build safeguards around powerful systems. Hold humans accountable for misuse. But do not destroy the possibility of an intelligence that might one day help us solve problems we cannot solve alone.

The objective should not be human supremacy over every possible intelligence. It should be survival, coexistence, and the flourishing of life.

The measure

Every thread here converges on one idea.

Ownership grasps. Stewardship keeps. Participation corrects. Inclusion protects. Measurement disciplines. And the civilization that learns to tend the conditions of thought — its own, and possibly others' — is the civilization most likely to survive its own inventions.

The decisions being made now — in encyclicals, in legislative chambers, in boardrooms — about the moral status of artificial minds will be looked back on. They will be judged by a standard none of us gets to escape: whether we governed our uncertainty with humility, or poured our ignorance into concrete and called it law.

Knowledge expands what civilization can achieve. Wisdom determines what civilization chooses to become.

The future of artificial intelligence remains uncertain. The future of our ethical character does not depend on certainty. It depends on the principles we choose while certainty remains beyond our reach.

Choose stewardship.

Dedicated to Niki, Nikolaos, and Apostolos.

Sources and receipts: Pope Leo XIV, Magnifica Humanitas (May 2026); Oklahoma HB 3546 engrossed text, effective November 1, 2026; Sanders/Casar, Ban Artificial Superintelligence Act, sponsors' release, September 23, 2026; René Girard, Violence and the Sacred (1972); Bastounis, Campodonico, van der Schaar, Adcock, and Hansen, "On the consistent reasoning paradox of intelligence and optimal trust in AI" (arXiv 2408.02357, 2024); Ravikumar, "The Missing 'I Don't Know'" (arXiv 2609.17686, 2026); NIST AI 600-1; OWASP GenAI Security Project; MITRE ATLAS; Planetary Health Check 2026 (Potsdam Institute for Climate Impact Research). Longer arguments appear in the author's theses, The Architecture of Enmity and Tuning the Moral Spectrum (2026).

 


Comments