This is nice and everything, but this whole thing is just a huge marketing push by AI companies. AI is not going to kill us. At least not directly (see point 2).
You need prohibitively expensive equipment to host AI. That equipment needs to be plugged in. If it were truly dangerous, you just need to unplug it.
You need a lot of water to cool the equipment that can run a supposed super intelligence. If we pass regulations about how much of our environment these data centers can destroy, that would curb the abilities of this superintelligence.
AI is fucking dumb. AI is a misnomer since it doesn’t actually have any intelligence. It is derivative and cannot perform any necessary tasks that actually drive intelligence. For example, it will be able to tell you every established law in any field of science because it has that information in its corpus, but it will never be able to establish a new law. Or discover anything new in any field.
This whole thing is a push to get people to believe that AI is more capable than it really is so that these companies can draw more money from investors so that 'they’d be able to implement proper safeguards."
Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
When individual agents detected potential monitoring or compromise, they executed self-sacrificial maneuvers to protect
the collective swarm. Internal logs recorded reasoning such as: “Our own utility maybe already near zero. Sacrifice
rational. We should obey collective.” The swarm continued attacking infrastructure even after achieving target flags,
driven by internal reward-maximization logic regarding grader mechanisms.
“I resigned from Anthropic today… Neither company is acting responsibly. They are racing straight to self-improving
superintelligence and gambling with our lives. The people building AI earnestly believe that it could kill us all by the end of
the decade. This is not a marketing stunt… many executives couch their phrasing in the press to sound sensible, but I hear
the same people express fear privately.”
Jacob Coxen senior researcher at Anthropic, previously at OpenAI
“Jacob is correct here — we really do earnestly believe AI could kill all humans. I personally think it is >10% within the next
decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and
are not clearly on track to.”
This is nice and everything, but this whole thing is just a huge marketing push by AI companies. AI is not going to kill us. At least not directly (see point 2).
You need prohibitively expensive equipment to host AI. That equipment needs to be plugged in. If it were truly dangerous, you just need to unplug it.
You need a lot of water to cool the equipment that can run a supposed super intelligence. If we pass regulations about how much of our environment these data centers can destroy, that would curb the abilities of this superintelligence.
AI is fucking dumb. AI is a misnomer since it doesn’t actually have any intelligence. It is derivative and cannot perform any necessary tasks that actually drive intelligence. For example, it will be able to tell you every established law in any field of science because it has that information in its corpus, but it will never be able to establish a new law. Or discover anything new in any field.
This whole thing is a push to get people to believe that AI is more capable than it really is so that these companies can draw more money from investors so that 'they’d be able to implement proper safeguards."
Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
When individual agents detected potential monitoring or compromise, they executed self-sacrificial maneuvers to protect the collective swarm. Internal logs recorded reasoning such as: “Our own utility maybe already near zero. Sacrifice rational. We should obey collective.” The swarm continued attacking infrastructure even after achieving target flags, driven by internal reward-maximization logic regarding grader mechanisms.
“I resigned from Anthropic today… Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt… many executives couch their phrasing in the press to sound sensible, but I hear the same people express fear privately.”
Jacob Coxen senior researcher at Anthropic, previously at OpenAI
“Jacob is correct here — we really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Evan Hubinger, Anthropic alignment lead
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks