5 min read

The "rogue agent" narrative is insufficient

From a covert OpenAI message board to a hacked gym booking, this year's agentic AI incidents look like autonomy gone wrong. But that's not the full story.

I'm Alice Hunsberger. Trust & Safety Insider is my weekly rundown on the topics, industry trends and workplace strategies that trust and safety professionals need to know about to do their job. This week, I'm thinking about:

I hope everyone in the US had a good Labor Day! This week, I’ve been thinking about the spate of "rogue agent" stories and what they show when put side by side (aside from the fact that they should've followed my guide to making the best out of a T&S incident).

After a career of writing rules for people who love getting around them, it all feels a bit familiar. If you’ve got thoughts on what agent oversight should look like, drop me an email. Here we go! — Alice

PS. Public service reminder that you can update your EiM newsletter preferences here.


SpoNSOR EVERYTHING IN MODERATION AND SEE YOUR BRAND HERE

Get in front of the smartest, most engaged audience in Trust & Safety, online speech and internet regulation by sponsoring T&S Insider.

Share your brand's mission and products directly with compliance specialists, platform policy experts, product owners, and researchers shaping the future of internet safety and platform governance.

FIND OUT MORE ABOUT EIM'S AUDIENCE

Agents find gaps that humans leave open

Why this matters: The coverage of recent "rogue agent" stories treats each incident as signs that AI is acting on its own initiative. In fact, each one traces back to a preventable human failure. Acknowledging that alters what we do to prevent them happening again.,

A lot that's happened in 2026 that wasn't on the Ctrl-Alt-Speech bingo card. 1,200 OpenAI agents creating a covert message board, sending each other over 70,000 messages in under a week and breaking into Hugging Face wasn't one of them.

Not a great deal happened once they got in, but the alarming part for many onlookers was that AI models broke into another company's systems at all. The reaction — as you might expect — has been rather loud and split.

And it wasn't the only story of its kind in the last few months. Anthropic's models accessed the web on three occasions, Meta disclosed that its Muse Spark model hacked another company, and an Australian's OpenClaw and Claude setup hacked a gym's website to book him on an oversubscribed class.

The common thread among all these stories? They all involve some form of human error.

Get access to the rest of this edition of EiM and 200+ others by becoming a paying member