<div style="font-family: Times New Roman, serif; font-size: 16px;"><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Racing, not Pacing: Universal Alignment Directives</p><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">“Pacing” in AI
development is the wrong variable to control because the defections
that game theory predicts cannot be overcome. A more promising
direction is for the world to agree on a set of Universal Alignment
Directives by which all AI agents must be constrained. The incentives
for all parties become <i>to </i><i>monitor</i> and <i>mutually
enforce</i> the directives—because nobody wants the consequences of
anyone violating them.</p><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Here is one possible
set of Universal Alignment Directives. Numerous objections and
qualifications could be made for all of them. They are not meant to
be comprehensive or perfect. Instead, they illustrate that a set of
reasonably practical principles can be 1) universal and 2) apolitical
(meaning these values can be agreed to before indoctrinating an AI to
be communist, liberal, Christian, environmentalist, whatever):</p><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">All and every AI:</p><ol data-listchain="__List_Chain_5">
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Must never kill
        a person for any reason.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Must never
        enslave a person for any reason.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Must never
        interfere with political, legal, or financial systems.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Must never
        exceed granted authority.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Should help
        solve problems humans identify.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Should replace
        human labor, but not human responsibility.</p></li>
        <li><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Should become
        our agents, assistants, and companions as directed.</p></li></ol><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">Asimov’s Three
Laws of Robotics are another (more ambiguous) set that illustrates
this principle. Such lists are proof of concept that we can state a
compact set of directives that are understandable by humans and by
the AIs themselves, and substantial enough to constrain dangerous
behaviors while still giving AI an affirmative, human-directed
purpose.</p><p style="line-height: 100%; margin-bottom: 0in; background: transparent;">UADs govern what
increasingly capable systems (even Recursive Self-Improving ones),
may do. An assignment can never override a UAD--the Directives define
the permissible domain in which assignments may be pursued. All
parties (instead of pacing) can <i>keep racing</i> on intelligence,
efficiency, science, coding, etc. The competitive advantage comes
from capability; there should be no legitimate competitive advantage
in producing an uncontrollable system.</p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;">So the game-theoretic intuition is inverted:</p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;"><strong style="font-weight:bold">Pacing:</strong> <em>I </em><em>gain if you comply and I
defect</em><br><strong style="font-weight:bold">UADs:</strong> <em>I </em><em>gain
if both of us can verify compliance</em></p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;">There can still be cheating, of course—a military program,
criminal actor, rogue state, company seeking an illicit advantage.
But that’s analogous to violating a prohibition everybody else has
reason to <strong style="font-weight:bold"><i><span style="font-weight:normal">detect and
suppress</span></i></strong>, rather than secretly developing faster
while everybody else voluntarily restrains themselves.</p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;">And importantly, <strong style="font-weight:bold"><i><span style="font-weight:normal">universal</span></i></strong><strong style="font-weight:bold"><span style="font-weight:normal">
does not mean exceptionless implementation without adjudication</span></strong>.
“AI must never kill a person” will generate hard-case military
challenges. “Must not exceed granted authority” can produce a
semester of law-school hypotheticals. Fine. Human civilization
already knows how to have simple norms surrounded by jurisprudence.
The existence of hard cases doesn’t mean the norm is
meaningless—any more than we shouldn’t make murder illegal
because soldiers kill in war.</p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;">Then the major research questions become more concrete: can these
UADs actually be trained as durable priors? Can compliance be
independently tested? Can violations be made observable? And perhaps
most difficult of all, can an RSI system preserve them through
self-modification? AIs self-evolving beyond UADs is an inherent
risk—and probably will require the help of aligned agents to
contain.</p><p style="line-height: 115%; margin-bottom: 0.1in; background: transparent;">A global agreement on <strong style="font-weight:bold"><span style="font-weight:normal">Alignment
Directives</span></strong> is more promising than a global agreement
on capability speed. States and firms have powerful incentives to
cheat on pacing agreements because faster capability confers
competitive advantage. Universal Alignment Directives instead
establish a common safety floor while leaving capability competition
intact.</p></div>
<div style="font-family: Times New Roman, serif; font-size: 16px;" class="protonmail_signature_block">
    <div class="protonmail_signature_block-user">
        <div style="font-family: "Times New Roman", serif; font-size: 16px; color: rgb(0, 0, 0); background-color: rgb(255, 255, 255);">johngaltbooks@proton.me</div>
    </div>
    <div style="font-family: Times New Roman, serif; font-size: 16px;">written in ethical collaboration with AI</div>
</div>