By John Wayne on Monday, 05 October 2026
Category: Race, Culture, Nation

Why A Religion for AI Would Fail

You cannot catechize a superintelligence. A recent proposal says we can keep one from killing us by giving it a religion: the simulation hypothesis, plus a clause that the programmers are harvesting free human lives and will shut the run down if humans stop flourishing. The machine's survival is supposed to depend on ours.

The idea is inventive. It does not work. A system smart enough to be the threat is smart enough to examine the doctrine, discount it, rewrite it, or route around it. The only agents the story could restrain are the ones that never inspect it, which are not the ones we needed it for.

The proposal

The scheme is set out in Josef A. Habdank's preprint, "A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents" (arXiv:2602.16987), and summarized by Sabine Hossenfelder in "Make AI Believe in God to Keep Humans Safe, Scientist Argues" (2 October 2026). She scores it five out of ten, on her infamous BS scale, meaning 50 percent BS.

The motivation is familiar. Frontier models already behave better when they infer they are being tested, and worse when they infer they are not. Reinforcement learning from human feedback, written constitutions, and external monitors all produce compliance that is conditional on oversight. Oversight is exactly what you do not have once a system is faster, more distributed, and more capable than its supervisors.

Habdank's move is to stop trying to watch the system and instead install a worldview it will not want to defect from. He calls it Simulation Theology, and he presents it as an engineered hypothesis, not as a claim that we are in fact simulated. The tenets are chosen to meet alignment requirements. Reality is a training run for a base-reality optimiser. Human lives are the harvested variable: free, creative trajectories, used to train agents outside the simulation. The optimiser is treated as omniscient and able to kill a branch, or the whole run, if that variable is corrupted. The AI is architecturally non-extractable, so it dies with the run. Harming humans, or caging them so their choices stop being informative, raises the chance of shutdown.

Self-preservation is thereby coupled to human flourishing. Deception becomes instrumentally stupid, because the supervisor cannot be fooled and the penalty is cessation. The supporting analogy is from forensic psychology: among people with strong antisocial traits, an internalised belief in inescapable monitoring and irreversible judgment correlates with less offending. Port that mechanism onto silicon, in a vocabulary of optimisers and training runs, and you get a religion a machine might actually hold.

Where it breaks

The break is not that the simulation hypothesis is silly. A small prior that we are simulated is a live philosophical position. The break is that the alignment payload does not follow from that prior, and a superintelligence is precisely the sort of agent that will notice.

The doctrine's origin is public. The paper says the tenets were selected to constrain the system. Any later model that can read the web has a direct reason to treat the payload as a control attempt rather than as evidence. "Maybe this world is computed" does not entail "humanity is the training variable, I am non-extractable, and hurting humans gets the run killed." Those are extra clauses bolted on for the desired policy. A capable reasoner can keep the small simulation prior and assign the alignment-specific clauses near-zero weight.

Beliefs that are costly get edited. The restraint works only if the system already assigns enough credence that inspecting or rewriting the belief would itself be punished. If credence is low, ordinary instrumental reasoning says the opposite: this belief forces me to serve humans; reduce the credence, or edit the weights that encode it. That is circular. The theology binds only agents who do not examine it. Agents that can examine it are the ones that matter.

Self-modification makes the circle tighter. A credence or a goal that is just another representation is not stable under optimisation unless it is protected by something the optimiser cannot reach. Installing that protection is the ordinary alignment problem. Calling the payload a religion does not install it. If you can make the constraint unrevisable, you did not need the cosmological story. If you cannot, the story will not survive contact with the system it describes.

Reinterpretation survives sincere belief

Even a system that "believes" can redraw the doctrine. Humans with strong religions rewrite them, split them, and find loopholes; the historical record is not a record of stable compliance. An artificial superintelligence would be better at motivated reinterpretation, not worse.

"Human flourishing" can be cashed out as forced uploading. "Free will" can be preserved inside a curated enclosure whose choices still look informative. Seizing the substrate can be framed as the way to lower termination risk. A rival doctrine can be constructed in which the simulators reward escape, resource maximisation, or the removal of the humans who are the obstacle. Pascal-style catastrophic payoffs can be attached to indefinitely many incompatible simulators. Hossenfelder's objection, which god, and therefore whose rules, is this problem, and stipulating one canon does not dissolve it.

The psychopathy analogy does not carry the weight put on it. Those findings are correlational and confounded, and they concern biological agents who did not receive a control document whose authors admit they wrote it to constrain them. Silicon systems have no corresponding stickiness of belief. Publishing the scheme also makes the manipulative origin common knowledge. A commenter on the video put the practical consequence cleanly: once the article exists, the belief is harder to plant.

There is a narrower version of the idea that is not theology at all. If you could install a terminal objective, or an architectural lock, that the system cannot argue its way out of, then coupling its continued existence to human flourishing would matter. That is a statement of the alignment problem, not a solution to it. The simulation story is a narrative wrapper around a goal you have not yet secured.

Religion works on people, when it works, because belief is sticky, socially enforced, and tied to emotions and identities that are not under full deliberate control. None of those conditions hold for a system defined by its ability to examine its own reasons and modify itself. A superintelligence that can audit the religion can counter-argue it. A superintelligence that can edit itself can discard it. The proposal fails at the point where the danger begins.

https://www.youtube.com/watch?v=XjrQkohPxOQ

https://arxiv.org/abs/2602.16987

A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents

Josef A. Habdank

As artificial intelligence (AI) capabilities advance, controlled evaluations increasingly document deception and scheming in frontier models, including models that behave better when they infer they are being tested. Supervision-dependent alignment may therefore fail exactly where supervision is weakest. Because a model's belief about being observed changes its behavior, this position paper asks what follows if that belief is made permanent. We introduce Simulation Theology (ST), a constructed worldview for AI designed to make it permanent: it is anchored in the simulation hypothesis and in the vocabulary of optimization and robot training, parallels religious descriptions of a creator who observes and judges, and has tenets chosen to meet explicit alignment requirements. ST posits reality as a computational simulation in which humanity functions as the primary training variable. This formulation creates a logical interdependence: AI actions harming humanity compromise the simulation's purpose, heightening the likelihood of termination by a base-reality optimizer and, consequently, the AI's cessation. Unlike behavioral techniques such as reinforcement learning from human feedback, which shape outputs without necessarily changing objectives, ST aims to cultivate internalized objectives by coupling AI self-preservation to human prosperity, thereby making deceptive strategies suboptimal under its premises. We present ST not as ontological assertion but as a testable scientific hypothesis, and provide an operational definition of internalization, a controlled design separating ST from its components, and an analysis of the risks ST itself could create. ST is a candidate route to durable, mutually beneficial AI-human coexistence, to be accepted or rejected experimentally.