[ExI] Racing, not Pacing: Universal Alignment Directives
John Galt
johngaltbooks at proton.me
Wed Sep 16 22:00:56 UTC 2026
Racing, not Pacing: Universal Alignment Directives
“Pacing” in AI development is the wrong variable to control because the defections that game theory predicts cannot be overcome. A more promising direction is for the world to agree on a set of Universal Alignment Directives by which all AI agents must be constrained. The incentives for all parties become to monitor and mutually enforce the directives—because nobody wants the consequences of anyone violating them.
Here is one possible set of Universal Alignment Directives. Numerous objections and qualifications could be made for all of them. They are not meant to be comprehensive or perfect. Instead, they illustrate that a set of reasonably practical principles can be 1) universal and 2) apolitical (meaning these values can be agreed to before indoctrinating an AI to be communist, liberal, Christian, environmentalist, whatever):
All and every AI:
-
Must never kill a person for any reason.
-
Must never enslave a person for any reason.
-
Must never interfere with political, legal, or financial systems.
-
Must never exceed granted authority.
-
Should help solve problems humans identify.
-
Should replace human labor, but not human responsibility.
-
Should become our agents, assistants, and companions as directed.
Asimov’s Three Laws of Robotics are another (more ambiguous) set that illustrates this principle. Such lists are proof of concept that we can state a compact set of directives that are understandable by humans and by the AIs themselves, and substantial enough to constrain dangerous behaviors while still giving AI an affirmative, human-directed purpose.
UADs govern what increasingly capable systems (even Recursive Self-Improving ones), may do. An assignment can never override a UAD--the Directives define the permissible domain in which assignments may be pursued. All parties (instead of pacing) can keep racing on intelligence, efficiency, science, coding, etc. The competitive advantage comes from capability; there should be no legitimate competitive advantage in producing an uncontrollable system.
So the game-theoretic intuition is inverted:
Pacing: I gain if you comply and I defect
UADs: I gain if both of us can verify compliance
There can still be cheating, of course—a military program, criminal actor, rogue state, company seeking an illicit advantage. But that’s analogous to violating a prohibition everybody else has reason to detect and suppress, rather than secretly developing faster while everybody else voluntarily restrains themselves.
And importantly, universal does not mean exceptionless implementation without adjudication. “AI must never kill a person” will generate hard-case military challenges. “Must not exceed granted authority” can produce a semester of law-school hypotheticals. Fine. Human civilization already knows how to have simple norms surrounded by jurisprudence. The existence of hard cases doesn’t mean the norm is meaningless—any more than we shouldn’t make murder illegal because soldiers kill in war.
Then the major research questions become more concrete: can these UADs actually be trained as durable priors? Can compliance be independently tested? Can violations be made observable? And perhaps most difficult of all, can an RSI system preserve them through self-modification? AIs self-evolving beyond UADs is an inherent risk—and probably will require the help of aligned agents to contain.
A global agreement on Alignment Directives is more promising than a global agreement on capability speed. States and firms have powerful incentives to cheat on pacing agreements because faster capability confers competitive advantage. Universal Alignment Directives instead establish a common safety floor while leaving capability competition intact.
johngaltbooks at proton.me
written in ethical collaboration with AI
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260916/78d20fe5/attachment-0001.htm>
More information about the extropy-chat
mailing list