[ExI] Racing, not Pacing: Universal Alignment Directives

Ben Zaiboc benzaiboc at proton.me
Fri Sep 18 06:57:50 UTC 2026


>>> All and every AI:
>>>
>>>     Must never kill a person for any reason.
>> We can stop here, because this is a clear one on which you definitely won't get agreement. I'm sure the US isn't the only nation in the world that would never agree to this. In fact, probably every nation on earth could come up with at least one reasonable exception to it.
>>
>> All such attempts to come up with a universal agreement, that everyone in the entire world has to really agree to (not just say they are, then ignore it), are doomed to failure. I don't see how that's not completely obvious.
> Ben, thank you for the very interesting reply. In every country murder is illegal. Yet every country also allows soldiers to kill under constrained circumstances. Nobody agreed to anything in order to have this consequence. Yes, militaries will use AI. The principle is what's important, not the exceptions. 
>
>> For problems like this, we could do worse than look to biology for a solution. That is, look at how biology solves these kinds of problems. The immediately obvious one is the continuously evolving battle between parasites and immune systems.
> Very interesting also, but say, corporate competition in a free(ish) marketplace resembles evolution like this only somewhat. We can and do have agreements and treaties in civilization. The point of my post was, if we have agreements, let's design ones that are more compatible with incentives rather than not.


So basically, you're saying that existing law should apply to agentic AI as well as to people.

Fair and reasonable.

The problem of various world states treating AIs, or at least some AIs, as special cases (like the military) will then arise. Note that existing law doesn't cover your 7 items for people, so wouldn't for AIs either.

I can see a case for effectively treating AIs the same as people (or at least AIs that are at least as intelligent as most people), and would agree with it, but that doesn't solve the alignment problem.

Or to look at it another way, how would you propose to apply your 7 items to humans (all humans)? If you could make that work, there may be a case for thinking it could work for AIs.

In fact, that would be a vast improvement to the general state of things, completely independent of the AI question.
-- 
Ben



More information about the extropy-chat mailing list