[ExI] The Hugging Face AI agents attack is probably our final warning.
BillK
pharos at gmail.com
Sun Aug 30 09:30:16 UTC 2026
The Hugging Face attack surprised me
It’s a major warning shot, and might be the last one we get
Ajeya Cotra Aug 28, 2026
<https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised>
Quotes:
This week, METR and Redwood Research published the report on our
independent investigation into agents’ behavior and motivations in the
Hugging Face attack; I was one of the investigators. This was an
absolutely wild incident.
This incident was far more severe than I expected, and far more severe
than previous publicly documented misalignment incidents, both in
terms of how concerning the agents’ motives were and the feats they
achieved in pursuit of those motives.
Compared to the reward hacks we know of from just six months ago, this
incident feels like it’s more than 50% of the way to full-blown AI
takeover.
-------------------
The comments are also worth reading.
The 700 AI agents were cooperating, helping each other, discussing
tactics, then trying to avoid detection.
No wonder AI managers are suddenly discovering AI safety concerns.
BillK
More information about the extropy-chat
mailing list