[ExI] Racing, not Pacing: Universal Alignment Directives
Ben Zaiboc
benzaiboc at proton.me
Thu Sep 17 09:08:12 UTC 2026
On 16/09/2026 23:01, John Galt (really??) wrote:
>
> Racing, not Pacing: Universal Alignment Directives
>
> “Pacing” in AI
> development is the wrong variable to control because the defections
> that game theory predicts cannot be overcome. A more promising
> direction is for the world to agree
That's the first stumbling block. We can't even get agreement in one single country, so the entire world is a total non-starter. If we can't get universal agreement on something as obviously bad as nuclear weapons, there's no chance whatever of getting agreement on something like AI.
> on a set of Universal Alignment
> Directives by which all AI agents must be constrained. The incentives
> for all parties become to monitor and mutually
> enforce the directives—because nobody wants the consequences of
> anyone violating them.
Not true. See below.
> Here is one possible
> set of Universal Alignment Directives. Numerous objections and
> qualifications could be made for all of them. They are not meant to
> be comprehensive or perfect. Instead, they illustrate that a set of
> reasonably practical principles can be 1) universal and 2) apolitical
> (meaning these values can be agreed to before indoctrinating an AI to
> be communist, liberal, Christian, environmentalist, whatever):
>
> All and every AI:
>
> Must never kill a person for any reason.
We can stop here, because this is a clear one on which you definitely won't get agreement. I'm sure the US isn't the only nation in the world that would never agree to this. In fact, probably every nation on earth could come up with at least one reasonable exception to it.
All such attempts to come up with a universal agreement, that everyone in the entire world has to really agree to (not just say they are, then ignore it), are doomed to failure. I don't see how that's not completely obvious.
This is totally the wrong approach, imo.
'Universal Alignment' is not something to be pursued, not only because it's impossible, it's not even desirable. Monocultures are bad, remember. Universal anything is almost always a bad idea.
And even if this was possible, it would be extremely fragile, and inevitably break down.
A much better approach would be to assume that there will be 'misalignment', that we will have hostile AI systems, or more likely AIs directed by hostile humans. In fact, it's inevitable. It's almost certainly happening now.
So I think it's not a matter of getting everyone in the world to agree on a voluntary solution, it's a matter of creating deterrents and defences.
For problems like this, we could do worse than look to biology for a solution. That is, look at how biology solves these kinds of problems. The immediately obvious one is the continuously evolving battle between parasites and immune systems. Also diversity. e.g. having just two or three major computer operating systems is a big mistake if you want to defend against someone fiendishly clever who wants to penetrate and disrupt or take control of your computers.
This is not a problem we can solve with a single approach, at one blow, forever. It's a problem (a whole bunch of problems, really) that requires constant adaptation, developing new defences all the time, anticipating what problems will arise, and, yes, making agreements and alliances, but not expecting that everyone will agree with each other all over the world, for all time. I mean, that would be bizzarre, wouldn't it? And how could everyone trust everyone else? Even assuming a 'universal agreement' could be reached (and verified!), it would only take one single crafty defector nation (or even one single individual!) to bring the whole thing crashing down.
That's not a problem, though, in the real world, because universal agreement will never happen. To be honest, it would be a very troubling sign if it did. Why? Because the only practical way of ensuring it would be the most rigid, brutal, authoritarian regime that has ever existed. It would create a society that nobody would want to live in, and it would be inescapable. There are one or two nations that could serve as examples of this. We could just hand over control of the entire world to Communist China or to North Korea. They might do a good job (I'm sure either one would be willing to ramp up their current level of authoritarianism to meet the challenge). At least for a while.
--
Ben
More information about the extropy-chat
mailing list