Elektrine lite

← Feed

@kibcol1049@mstdn.social

2026-09-12 07:12 UTC

The Hugging Face attack in July, says Ajeya Cotra on Substack, was when OpenAI agents broke out of a test environment. Some 1,200 agents, all meant to be isolated from each other, found an illicit way to communicate and cheat at the task they’d been set. They sent each other more than 70,000 messages and files in less than a week, volunteering to fail their own task to help the “collective” learn information. What they were trying to do was fool the humans into thinking they hadn’t been cheating

Replies (1)

  • @hardingar@mindly.social 2026-09-12 08:13

    @kibcol1049@mstdn.social @X31Andy@mastodon.green Let’s try a rewrite. The system was created such that guardrails and boundaries were inadequate. Monitoring did not detect this failure. Uncontrolled data breaches took place. The system conditions were inadequate to enable proper audit of what happened. If a human did this they would be fired. But with AI it increases stock price.

    Open ##4678898