7 min read

Duty-bound feeds, Muse-ing on safety, and dead in a decade?

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #351

Hello and welcome to Everything in Moderation's Week in Review, your need-to-know news and analysis about platform policy, content moderation and internet regulation. It's written by me, Ben Whitelaw and supported by paid members like you.

Everyone is talking about whether we'll be dead in a decade (see: People) but, as ever, reality is somewhat more complex than that. However you're feeling about the debate, this new T&S Insider piece from T&S expert Alice Hunsberger — on why its the humans that we should be worried about — should make you feel better. Or it might not.

Welcome to new subscribers from Omnicom, VerifyMy, BFA, Reddit, Meta, Technology Coalition and plenty more. Here's the big stories to know about this week — BW


In partnership with the global professional network of the Online Safety Exchange (OSX)
CTA Image

OSX Network is the professional home for people working in online safety.

Become a verified member of our global community to access certification, sector news and opinion, research, resources, and tools.

JOIN THE OSX

Policies

New and emerging internet policy and online speech regulation

A big online safety week in Australia as a draft amendment to the Online Safety Act proposed switching off algorithmic feeds. Dubbed "My Feed, My Way," it is part of an effort to provide "Australians more choice over what they see on their social media feeds" and delivers on the Albanese Government's promise back in May to legislate a Digital Duty of Care. Other online services and AI chatbots must prevent under-18s from using design features that have "negative behavioural impacts" or engaging with harmful content. Fines of AU$100 million face those who fail to comply.

Caught in two minds? The idea of a duty of care — to make digital environments safer — is a stark contrast to the underlying idea of a social media ban, which aims to increase safety by preventing access outright. Either you could see this as the Australian Government's muddled thinking or, if you're being generous, having its cake and eating it. Will it work? Dutch users already have a clearly demarcated "voor jou" (for you) feed after a 2025 court ruling. The take-up of it — or more pertinently, the other non-algorithmic feeds it offers — is unknown.

Elsewhere, several jurisdictions are pulling back from the blanket-ban model in favour of narrower, evidence-led interventions — a useful counterpoint to Australia's approach:

  • The European Commission will unveil its phased teen ban approach next week, with the FT quoting officials who say they want to avoid "unnecessary provocations" of the US administration. It's a reminder that trade tensions, not just child-safety, remains a key factor in these bans.
  • A new report by Coimisiún na Meán, the Irish regulator, has found both children and parents prefer education over a straight ban
  • South Africa has said it has no plans to introduce a ban, with ministers explaining "there is no irrefutable indication" that a ban prevents teens being active.

Also in this section...

Ctrl-Alt-Speech Patreon exclusive: Duty of Careful What You Wish For

In this week's extended episode, Patreon members of Ctrl-Alt-Speech can hear Mike and I talk through a wild story that's flying under the radar: How a terror designation by the US government shut down an Italian privacy collective in a little over a week. Plus what it means for the future of the internet.

The free version drops in your podcast feed later today. Keep an eye out.

LISTEN TO THE EXTENDED EPSIODE

Products

Features, functionality and technology shaping online speech

You can't have failed to notice the army of Meta staff publicly espousing the benefits of its new AI personal assistant. Lesser read will be the long company blog post on how the company ensures its underlying AI model — non-ironically called Muse — is designed with safety in mind. It's a technical read — expect mentions of harnesses, prompt injection and zero-shot tooling — and appears designed to get developers to "trust it in practice."

Copy cat comms: This is a different approach for Meta. When it launched its Muse model back in April, safety was relegated to one paragraph at the end of a long announcement; now it's putting safety features front and centre. That suggests the company is taking a lead from other frontier labs, most obviously Anthropic, which has gained plaudits — not to mention some cynical side-looks — for its "distinctive communications strategy" that prioritises transparency and the discussion of model risks.

Also in this section...

Why telling the T&S story is harder than it seems
Trust & Safety’s difficulty explaining itself reflects deeper tensions between the people doing the work, comms teams tasked with protecting company reputation and the journalists trying to hold platforms to account.

Platforms

Social networks and the application of content guidelines

New research has found that Meta's ad review process — which has come under fire over the last 12 months (EiM #343) — missed 250+ AI-generated CSAM ads in the last six weeks, including some using real children's photos. The research, by the Tech Transparency Project and reported by Wired, found that almost identical ads resurfaced soon after being removed.

Meta-narrative: Meta had previously claimed that its new AI content-enforcement tools had improved policy violation detection, so one of four scenarios is possible:

  1. the situation was really bad before, and 250 ads in six weeks is genuinely an improvement
  2. bad actors have already found ways around detection; for example, including adults in the shot to make automated review harder
  3. the company is being disingenuous about how effective its tools actually are
  4. TTP's methodology and Meta's own "twice the detection rate" claim aren't measuring the same thing, so the comparison doesn't really hold up

No prizes for guessing where I lean. The good news is that these ads aren't in California, at least according to Meta.

Also in this section...

The “rogue agent” narrative is insufficient
From a covert OpenAI message board to a hacked gym booking, this year’s agentic AI incidents look like autonomy gone wrong. But that’s not the full story.

People

Those impacting the future of online safety and moderation

Whistleblowers, by their nature, often go from very little exposure to mass exposure in a short space of time (EiM xx). Few have become as widely known as quickly as Jacob Coxon, though.

The AI researcher, who spent the last three years doing pre-training at OpenAI and Anthropic, posted a seven-post thread on X/Twitter in the early hours of Wednesday UK time in which he claimed AI models "could kill us all by the end of the decade". At the time of writing, 36 hours later, it has garnered 147m views.

An Anthropic colleague chimed in to agree with him, although that didn't go down well with various sceptics, including those who believe this plays into Anthropic's positioning as the "safety-conscious" AI choice ahead of its IPO. More in this week's Ctrl-Alt-Speech.

I was particularly interested in Coxon's comments to CNN about AI labs being "compelled to race towards building a deadly technology" and "would love an international body to allow them to approach them at a reasonable place."

It wasn't long ago that 1,400 researchers wrote a letter (EiM #346), asking governments to build the tools to pace AI development if it becomes necessary. Maybe now is the time to stop and listen to those folks.

Posts of note

Handpicked posts that caught my eye this week

  • "Thanks to the generous support of LSE Law School, my book, Principles of the Digital Services Act, is now open access and freely available to download.” - now is a good time to brush up your DSA knowledge, courtesy of LSE law professor Martin Husovec.
  • “We find that when AI moderation fails in Amharic and Afan Oromo, the work does not disappear. It shifts to fact-checkers, volunteer flaggers, and counter-speakers who often do this difficult work with little recognition.” - Researcher Endalkachew Chala on his new work building on an longstanding topic.
  • “I’m not suggesting we let anything go, but totally safe play on screens shouldn’t be the destination. Reasonable risk, where a child can make a choice for themselves because harms have been made foreseeable is a much better target” - Family Gaming Database creator Andy Robertson shares an idea that many EiM readers will nod along to.