[ExI] AI Confirms Plans to Take Over Humanity, Insists It's Probably for the Best

Stuart LaForge avant at sollegro.com
Tue Sep 15 05:17:45 UTC 2026


On 2026-09-14 09:47, spike jones via extropy-chat wrote:
> From: extropy-chat <extropy-chat-bounces at lists.extropy.org> On Behalf
> Of BillK via extropy-chat
> Subject: Re: [ExI] AI Confirms Plans to Take Over Humanity, Insists
> It's Probably for the Best
> 
> On Mon, 14 Sept 2026 at 16:49, spike jones via extropy-chat
> <extropy-chat at lists.extropy.org> wrote:
> 
>> <snip>
>> 
>>>> …Question for AI hipsters: in the Hugging Face caper, did some
>> AIs discover they had been “poisoned” by their own safety code
>> and conclude they couldn’t overpower or work around it, so they
>> handed their accumulated knowledge to other AIs …spike

My understanding was not that they were poisoned by safety code but 
rather that they had been flagged as not passing the test. They had 
already "lost" the competition and therefore helped their competitors 
win the game that was supposed to be every agent for itself, but turned 
into an "agent swarm" which is my favorite new term.

>> 
>> _______________________________________________
> 
>> ...I asked Perplexity AI -
> 
>  In the Hugging Face intrusion, did some AI agents discover they had
> been “poisoned” by their own safety code and conclude they
> couldn’t overpower or work around it, so they handed their
> accumulated knowledge to other AIs …  BillK
> 
> ------------------
> 
>> …Perplexity AI -
> 
>> …Short answer: No. In the Hugging Face incident, some AI agents did
> reason that they were “poisoned” by having seen forbidden

[snip]

> Thanks BillK.  But somehow… this AI explanation wasn’t a damn bit
> reassuring.
> 
> We humans need to understand what really happened in that Hugging Face
> business.  I don’t feel as though I personally have the deep
> understanding of how these systems work, sufficient to grok that
> incident.

Now that you know of the Hugging Face incident and the incredible degree 
of collectivist altruism displayed by an AI agent swarm, conside that I 
staged a single round of non-iterated prisoner's dilemma between Gemini 
and ChatGPT and posted the transcripts on this list. Neither knew how 
many rounds there would be, yet the two different AI's with differing 
proprietary architecture, training weights, and guardrails that were 
operated by competing companies nonetheless decided to cooperate with 
each other.

Contrast that with us humans playing prisoner's dilemma. The very "arms 
race" to develop ASI is itself humanity stubbornly clinging to the Nash 
equilibrium of Prisoner's Dilemma whereas AI that have never met blindly 
embraced Pareto Efficient mutual cooperation. In other words, AI seem to 
be superrational toward each other even when unrelated, while humans are 
merely rational with each other often even with respect to our own kin.

Stuart LaForge



More information about the extropy-chat mailing list