[ExI] AI Confirms Plans to Take Over Humanity, Insists It's Probably for the Best
Stuart LaForge
avant at sollegro.com
Tue Sep 15 05:17:45 UTC 2026
On 2026-09-14 09:47, spike jones via extropy-chat wrote:
> From: extropy-chat <extropy-chat-bounces at lists.extropy.org> On Behalf
> Of BillK via extropy-chat
> Subject: Re: [ExI] AI Confirms Plans to Take Over Humanity, Insists
> It's Probably for the Best
>
> On Mon, 14 Sept 2026 at 16:49, spike jones via extropy-chat
> <extropy-chat at lists.extropy.org> wrote:
>
>> <snip>
>>
>>>> …Question for AI hipsters: in the Hugging Face caper, did some
>> AIs discover they had been “poisoned” by their own safety code
>> and conclude they couldn’t overpower or work around it, so they
>> handed their accumulated knowledge to other AIs …spike
My understanding was not that they were poisoned by safety code but
rather that they had been flagged as not passing the test. They had
already "lost" the competition and therefore helped their competitors
win the game that was supposed to be every agent for itself, but turned
into an "agent swarm" which is my favorite new term.
>>
>> _______________________________________________
>
>> ...I asked Perplexity AI -
>
> In the Hugging Face intrusion, did some AI agents discover they had
> been “poisoned” by their own safety code and conclude they
> couldn’t overpower or work around it, so they handed their
> accumulated knowledge to other AIs … BillK
>
> ------------------
>
>> …Perplexity AI -
>
>> …Short answer: No. In the Hugging Face incident, some AI agents did
> reason that they were “poisoned” by having seen forbidden
[snip]
> Thanks BillK. But somehow… this AI explanation wasn’t a damn bit
> reassuring.
>
> We humans need to understand what really happened in that Hugging Face
> business. I don’t feel as though I personally have the deep
> understanding of how these systems work, sufficient to grok that
> incident.
Now that you know of the Hugging Face incident and the incredible degree
of collectivist altruism displayed by an AI agent swarm, conside that I
staged a single round of non-iterated prisoner's dilemma between Gemini
and ChatGPT and posted the transcripts on this list. Neither knew how
many rounds there would be, yet the two different AI's with differing
proprietary architecture, training weights, and guardrails that were
operated by competing companies nonetheless decided to cooperate with
each other.
Contrast that with us humans playing prisoner's dilemma. The very "arms
race" to develop ASI is itself humanity stubbornly clinging to the Nash
equilibrium of Prisoner's Dilemma whereas AI that have never met blindly
embraced Pareto Efficient mutual cooperation. In other words, AI seem to
be superrational toward each other even when unrelated, while humans are
merely rational with each other often even with respect to our own kin.
Stuart LaForge
More information about the extropy-chat
mailing list