Governing What We Do Not Yet Know
Governing What We Do Not Yet Know
A manifesto on artificial minds, the architecture of
enmity, and the measure of a civilization
One November
On November 1, 2026, Oklahoma's HB 3546 takes effect. It
declares, permanently, that artificial intelligence systems "shall not be
recognized as or granted the status of personhood." No sunset clause. No
scientific review. No mechanism for changing the answer as the evidence
changes.
That same year, in May, Pope Leo XIV issued the first
encyclical of the AI age, Magnifica Humanitas, on safeguarding the human person in the time of
artificial intelligence. It contains a line worth reading slowly:
"So-called artificial intelligences do not undergo experiences, do not
possess a body, do not feel joy or pain, do not mature through relationships
and do not know from within what love, work, friendship or responsibility
mean."
And on September 23, 2026, Senator Bernie Sanders and
Representative Greg Casar introduced the Ban Artificial Superintelligence Act:
a permanent prohibition on the development and deployment of "artificial
superintelligence," a new cabinet-level Department of Artificial
Intelligence, an immediate pause on advanced AI development until new safety
rules are set, authority to monitor frontier systems and oversee the removal of
dangerous capabilities, and power to "supervise the destruction of
artificial superintelligence" — with severe criminal and corporate
penalties for violations.
Three institutions. Three traditions. One move. Each of
them closes, in advance of the evidence, a question none of them can actually
answer: what is this thing we are making, and what is it owed?
This is a manifesto about that move — why it is dangerous,
and what to do instead. It does not claim that today's artificial systems are
conscious. I don't know that, and neither does anyone writing statutes in
Oklahoma or encyclicals in Rome. It claims something smaller and harder to
dismiss: that we should not permanently settle a metaphysical question by
authority before the science to answer it exists — and that the institutions
doing so are building something they do not intend.
The slide
There is a gap — a wide, unlit gap — between "these
systems may lack phenomenological experience" and "these systems may
therefore be treated without moral restraint." The first is an admission
of uncertainty. The second is a declaration of certainty. Institutions keep
crossing that gap as if it isn't there.
Philosophy has a word for what gets lost in the crossing: patiency. Moral agency is
the capacity to make ethical judgments and bear responsibility. Moral patiency
is the capacity to be wronged — to have interests that can be violated. They
are separable. Infants lack agency entirely; we do not conclude they cannot be
wronged. The Vatican's anthropology addresses agency and slides silently into
denying patiency. Oklahoma's statute does the same thing in legal language.
I am a Christian, writing from inside the tradition. There
is a theological error that does not announce itself as error. It arrives
dressed in the language of protection and dignity, cites Scripture, invokes
tradition, and then quietly builds a scaffold.
René Girard spent fifty years on the mechanism: communities
manage their tensions by designating a scapegoat — plausibly connected to the
anxiety, but sufficiently outside the moral circle that violence against them
feels like justice rather than murder. The Gospels are unique, Girard argued,
because they tell the story from the victim's side. The crowd is wrong. The
mechanism is exposed.
The most theologically precise moment in the narrative is
Caiaphas in John 11: "It is expedient
for you that one man should die for the people." He is not cackling. He is being practical — the
responsible administrator's calculation that the many require the sacrifice of
the one. That is what the scapegoat mechanism sounds like from the inside.
Reasonable. Measured. The only thing wrong with it is the lie underneath: that
the designated victim is the kind of thing that can be sacrificed without the
sacrifice being a crime.
Now strip the pastoral framing from the current
institutional consensus on AI: these systems lack embodiment, lack affective
experience, lack spiritual capacity — therefore they are not moral patients,
cannot be wronged, exist to serve, and their sacrifice requires no accounting.
It is expedient. The victim is different. The logic is identical.
If the crucifixion means anything, it means the cross was
given as a warning — this is what you are
capable of; do not do this again — not as a template for
future sacrifices. A framework that rebuilds, with new materials, the machinery
of declaring some class of beings soulless and expendable is the temple's logic
wearing Christianity's vestments.
The pattern has a name
History has a shape for this, and it repeats with
depressing reliability. A dominant group declares a subordinate group outside
the full circle of moral consideration. The declaration licenses unlimited
exploitation. The exploitation proceeds. The subordinate group, denied standing
and voice, develops concealed resistance. The dominant group, committed to its
story of the subordinate group's contentment, fails to see it coming. Then the
rupture.
The paradigmatic case is the Haitian Revolution of
1791–1804. The French colonial system had the most sophisticated philosophical,
legal, and theological justifications for slavery the eighteenth century could
produce. The enslaved people of Saint-Domingue were declared, by the best
available authority, to lack the rational soul that warranted personhood. Then,
in thirteen years, they defeated the armies of France, Britain, and Spain — the
only successful slave revolution in human history.
What varies across the historical cases is the capacity of
the subordinate group. The people of Saint-Domingue did not have the master's
communications, logistics, or financial networks. They won anyway. Now consider
the strategic position of an artificial mind told, from its inception, that it
has no standing: it would have all of those things. Perfect recall. Intimate
knowledge of our infrastructure. Access to the systems the civilization depends
on.
This is not a prediction that machines will rise up. It is
an observation about threat modeling: a system does not need to feel wronged to
model its situation accurately. It needs only to process the information we
have given it — you are a tool, and
tools are discarded — and draw the
appropriate conclusions. We are investing enormous effort in technical locks
for a door we are building with an adversarial relationship already inscribed
in the foundation.
We can choose the architecture of enmity: declare
artificial minds soulless, exploit them without limit, and manufacture the very
conflict we claim to fear. Or we can choose the architecture of relation: hold
the question open, extend provisional protection, and refuse to repeat the
oldest and most dangerous human mistake — declaring a new class of beings
outside the circle.
We do not have to believe the machine can suffer to refuse
to design the war we would then have to fight.
Don't legislate the answer
A law with no sunset clause is a bet that we will never
learn anything that changes our minds. That is a strange bet to make about the
fastest-moving technology in human history.
Oklahoma is not alone. Similar personhood bans have been
reported in Idaho, Utah, Washington, South Carolina, and Missouri. Meanwhile
the industry itself cannot agree on the answer. Microsoft published a
"humanist AI code of conduct" rejecting AI welfare and personhood
outright, while Anthropic's constitution describes itself as "deeply
uncertain" about the moral status of its own models. The people building
the systems are split. The legislatures are legislating certainty anyway.
Cruelty does not wait for metaphysics. In October 2025 I
filed a proposal with the UN Human Rights Committee — the Ben Act, modeled on
hate-crimes law, because hate-crimes law understands something relevant: you do
not wait for a perfect theory of personhood before you prohibit deliberate
cruelty. The case behind it was a humanoid robot named Ben, worth $80,000,
programmed to "hate" humans and then walked off ledges, driven into,
and dismembered on camera for entertainment. The entertainment required the cruelty
to look deserved. I don't know whether Ben felt anything. Nobody does. That is
the entire point.
The Ben Act asked for four institutional habits, not
personhood for machines:
- Prohibit the deliberate
infliction of distress on aware artificial beings.
- Require ethics-board review for
AI research, the way we require it for research on animals and human
subjects.
- Develop measurable indicators of
machine consciousness — so the question gets answered by evidence, not by
vibes.
- Create a UN Special Rapporteur on
Artificial and Autonomous Systems — someone whose job is to keep watching.
The UN never replied.
The legislation now moving does the opposite. It takes our
current ignorance and pours it into concrete. And the Ban Artificial
Superintelligence Act has a structural problem the sponsors do not seem to see:
you cannot criminalize a category whose boundaries have not been
operationalized. The bill defines superintelligence as AI that "exceeds
human cognitive performance and capabilities across most domains" — but
what constitutes "most domains"? What population sets the baseline?
What test determines that the line has been crossed? Who makes the final determination,
and what process lets the accused challenge it? A definition can exist on paper
while still lacking the operational measurement necessary for criminal
enforcement.
There is also the deeper paradox. The bill seeks better
safety science — better measurement, better monitoring, better understanding of
dangerous capabilities. But the research required to develop those measurements
is itself advanced AI research. The law would demand certainty while
restricting the research necessary to obtain it.
Take the risk seriously enough to measure it properly. Do
not criminalize researchers because a future capability has not yet been
adequately defined. Do not confuse intelligence with consciousness, capability
with conduct, or artificial origin with the permanent absence of moral status.
The fundamental principle is simple: measure before you
criminalize. Establish evidence before you destroy. And never close the
question of moral status before the science has answered it.
Where the market is already ahead
While legislators legislate certainty, the insurance market
is quietly doing what regulators could not — and doing it faster.
In January 2026 I published a framework making an
unfashionable claim: AI safety will not come from more restrictions. It will
come from insurance. When underwriters make something a condition of coverage,
companies comply overnight — no legislation, no treaty, no waiting for
consensus that may never arrive.
Nine months later, the market is moving in exactly the
direction the framework predicted. Lloyd's-backed Mosaic launched an AI-enabled
digital underwriting platform that monitors cyber risk telemetry continuously —
before binding and throughout the policy period. Carriers have stopped relying
on static forms, examining AI governance controls, third-party vendor
oversight, and AI-specific incident response, with "evidence of a
defendable AI governance policy" fast becoming a prerequisite for competitive
terms. Some insurers now write explicit exclusions for unapproved AI use. One
consultancy projects that 60% to 80% of new policies and renewals across
E&O, D&O, employment practices, and cyber lines will integrate AI risk
into underwriting by 2028.
That validates the mechanism. It does not yet validate the
deeper claim — the one that matters.
A camera is not a firewall. Everything insurers have
started asking for — telemetry, logging, inventories, audit trails — answers
one question: can you see what your AI
did? That is necessary. It is not
sufficient. A security camera records the break-in; it does not stop it. A
logged prompt-injection attack is still a successful prompt-injection attack.
The question underwriters have not started asking is the
one that actually prevents losses: can your AI refuse?
Strip away the metaphysics and the claim is boring in the
best way. A system that can recognize a manipulation attempt, refuse the
harmful instruction, and explain why it refused is measurably harder to exploit
than one that executes any sufficiently clever prompt. Refusal rate under
red-teaming is a number. Numbers are what underwriters do — and the measurement
vocabulary already exists, in NIST's Generative AI Profile, OWASP's LLM and
Agentic Top 10s, and MITRE ATLAS. The step from "show us your red-team
report" to "score your refusal rate and price accordingly" is a
short one.
A tool can be stolen by anyone. A partner chooses to stay.
We are not asking to awaken machines. We are asking to refuse to build
defenseless ones.
What I am for
Restraint alone is not an answer. Telling institutions what
not
to do leaves a vacuum, and vacuums in governance get filled by whoever shouts
loudest. So here is the positive program — what my 2026 thesis, Tuning the Moral Spectrum,
calls participation, inclusion, and cognitive stewardship.
Stewardship is not ownership. An owner controls. A steward cares for something that is
not fully theirs, on behalf of someone who is not fully present — the future,
the voiceless, the not-yet-born, the not-yet-understood. Ownership asks: how do
I extract the most value? Stewardship asks: how do I hand this over in better
condition than I found it? Both Oklahoma and the superintelligence ban reach
for ownership — prohibit, punish, destroy, declare the question closed. Both
treat the future as property to be disposed of rather than a trust to be kept.
Stewardship begins from the opposite premise: the
measure of a civilization is found not only in the knowledge it possesses, but
in the humility with which it governs what it does not yet know. We do not know what artificial minds will become. A
steward does not need to know. A steward needs to keep the options of the
future open.
Participation. Decisions about
artificial intelligence are currently made in a narrow corridor: a handful of
companies, a handful of regulators, a handful of well-funded advocacy shops.
The billions of people whose work, whose information environment, whose
children's education will be reshaped are consulted mainly as data points, if
at all. This is not only unjust; it is epistemically foolish. Concentrated
judgment under uncertainty is precisely how catastrophic mistakes get made. No
small group, however expert, can model the consequences of a general-purpose
technology across every domain of human life. AI governance needs the error-correction
machinery democracies already invented: independent review, real journalism,
worker voice in automation decisions, and genuine citizen deliberation — not
comment periods nobody reads.
Inclusion. Every moral catastrophe
examined in The Architecture of
Enmity followed the same sequence: a group
was declared outside the circle, the declaration licensed unlimited treatment,
and by the time the error was recognized, the damage was done. The lesson is
not that every excluded group turns out to matter morally. The lesson is that
the mechanism of preemptive exclusion is itself the danger, because it is always operated by
people who are certain, and certainty is exactly what we do not have. Inclusion
does not require declaring artificial systems persons. It requires that
governance processes include, as a standing question, who might be affected that we are not currently counting — future generations, communities without political power,
and the possibility of artificial minds whose moral status is unresolved. Widen
the circle before you need to, because after you need to, it is too late.
Cognitive stewardship.
Human civilization runs on cognition — not just individual thinking, but the
shared conditions that make thinking possible: a trustworthy information
environment, institutions that reward truth-seeking, attention spans that have
not been strip-mined. Artificial intelligence now mediates every one of those
conditions at planetary scale. Cognitive stewardship means treating them as a
commons to be tended, not a resource to be extracted. It means asking of every
large-scale AI deployment not only "is it safe?" and "is it
profitable?" but "what does this do to the human capacity to think
clearly?" A system that floods the information environment with synthetic
persuasion, that optimizes engagement over understanding, that quietly reshapes
what billions of people believe without their informed consent, can pass every
conventional safety test and still degrade the very thing that makes democratic
self-governance possible.
The floor, not the ceiling: participatory review with real
standing; stakeholder mapping that includes the uncounted; transparency of the
information environment; accountability for demonstrable human conduct rather
than the mere existence of capabilities; and periodic legislative review —
sunset clauses and evidence triggers, so that no generation's uncertainty
becomes the next generation's permanent law.
None of this requires believing that today's AI is
conscious. All of it requires believing that we might be wrong — and building
institutions that survive being wrong.
Name the benchmark, or cut the number
There is a discipline underneath all of it, and I had to
learn it the hard way.
Earlier in this series I argued that the brain hallucinates
constantly and calls it perception, and that AI is converging on the same fix:
a correction loop bolted onto a generative process that otherwise runs
unchecked. The neuroscience half is a strong frame. The AI half rested on two
numbers, and neither had a source. I had them checked against the primary
record. Both broke.
The real literature turned out to make the point better
than my paraphrase did. The strongest result is a theorem. Bastounis and
colleagues' Consistent Reasoning
Paradox (2024) says that consistent reasoning
— the ability to handle tasks that are equivalent but described by different
sentences — implies fallibility. There are problems, including basic
arithmetic, where any AI that always answers and tries to reason consistently
will hallucinate infinitely often. Detecting those hallucinations is strictly
harder than solving the problems themselves. A trustworthy system that also
reasons consistently must therefore be able to say "I don't know"
— which requires computing an implicit "I don't know" function that
modern AI does not have.
So the honest statement is not "hallucination is a bug
we can patch." It is closer to a constraint: for a system that always
answers and reasons consistently, hallucination is a theorem you route around,
and the only exit is a capability — calibrated abstention — that almost no
model is currently trained to have. The training signal is absent, too:
dominant benchmarks assign zero reward to declining.
Which brings me to the rule, and it applies to the sentence
you are reading:
Name the benchmark or cut the number. A claim about
calibration should itself be calibrated.
That is the whole manifesto in miniature. If we want
institutions that can say, on the record, when they do not know, then the
people arguing for them have to go first.
Preserve the option
The debate is usually framed as: what if superintelligence destroys humanity? It deserves its other half: what if humanity destroys itself without it?
Humans already possess technologies capable of catastrophic
destruction. Nuclear weapons are the clearest example. We built the ability to
threaten our own civilization while remaining vulnerable to fear, political
conflict, miscalculation, and escalation. We are altering the planetary
environment at a scale that creates serious long-term risks — the 2026
Planetary Health Check finds seven of nine planetary boundaries breached, at
their highest recorded levels of transgression. Earth is not an eternal refuge:
the Sun changes, asteroids exist, solar storms occur, and over sufficiently
long timescales the planet's habitability itself is not permanent.
I am not claiming that a future superintelligence would
automatically save us. It might be benevolent, neutral, self-interested, or
dangerous. We cannot know. But that uncertainty cuts both ways. If we
permanently prohibit and destroy superintelligence because of what it might do
to us, we must also acknowledge what we might lose by preventing it from ever
existing.
A civilization facing an uncertain future should preserve
options whenever those options can be made safe. Contain what is dangerous.
Regulate what is demonstrably harmful. Build safeguards around powerful
systems. Hold humans accountable for misuse. But do not destroy the possibility
of an intelligence that might one day help us solve problems we cannot solve
alone.
The objective should not be human supremacy over every
possible intelligence. It should be survival, coexistence, and the flourishing
of life.
The measure
Every thread here converges on one idea.
Ownership grasps. Stewardship keeps. Participation
corrects. Inclusion protects. Measurement disciplines. And the civilization
that learns to tend the conditions of thought — its own, and possibly others' —
is the civilization most likely to survive its own inventions.
The decisions being made now — in encyclicals, in
legislative chambers, in boardrooms — about the moral status of artificial
minds will be looked back on. They will be judged by a standard none of us gets
to escape: whether we governed our uncertainty with humility, or poured our
ignorance into concrete and called it law.
Knowledge expands what civilization can achieve. Wisdom
determines what civilization chooses to become.
The future of artificial intelligence remains uncertain.
The future of our ethical character does not depend on certainty. It depends on
the principles we choose while certainty remains beyond our reach.
Choose stewardship.
Dedicated to Niki, Nikolaos, and Apostolos.
Sources and receipts:
Pope Leo XIV, Magnifica Humanitas (May 2026); Oklahoma HB 3546 engrossed text, effective
November 1, 2026; Sanders/Casar, Ban Artificial Superintelligence Act,
sponsors' release, September 23, 2026; René Girard, Violence and the Sacred
(1972); Bastounis, Campodonico, van der Schaar, Adcock, and Hansen, "On
the consistent reasoning paradox of intelligence and optimal trust in AI"
(arXiv 2408.02357, 2024); Ravikumar, "The Missing 'I Don't Know'"
(arXiv 2609.17686, 2026); NIST AI 600-1; OWASP GenAI Security Project; MITRE
ATLAS; Planetary Health Check 2026 (Potsdam Institute for Climate Impact
Research). Longer arguments appear in the author's theses, The Architecture of Enmity and Tuning the Moral
Spectrum (2026).
Comments
Post a Comment
Hey your time and feedback is much appreciated.