Watermarked words, OpenAI's very ethical exit another Reddit revolt?
Hello and welcome to Everything in Moderation's Week in Review, your need-to-know news and analysis about platform policy, content moderation and internet regulation. It's written by me, Ben Whitelaw and supported by paid members like you.
I watched this week's solar eclipse with neighbours and in awe, only to come back online to find recycled footage and dodgy Photoshops among the many glorious, real photos of this amazing natural phenomenon. It's a reminder that everything is an information-quality event nowadays. The residents of Ceuta (See: Policies and this week's Ctrl-Alt-Speech) know that all too well.
Welcome to new free subscribers from Yoti, Vestiaire Collective, Ofcom, Stanford University, Christchurch Call, Google, Thorn and elsewhere plus the readers who have upgraded to paid membership in recent weeks. EiM only works because enough of you think independent, specialist coverage worth supporting.
I’m away on holiday next week so there’ll be no regular Week in Review. To be frank, I’m not a big fan of — or very good at — taking weeks off, and it’s notoriously hard for independent media to do so, as explained here. But it's necessary and important so thanks for bearing with me while I go splash around in the sea with my family.
Here's all you need to know from the last seven days — BW
New feature! Context Labels add more information to suspected CSAM.
As part of Thorn’s continued commitment to helping trust and safety teams detect and disrupt child sexual abuse material (CSAM), we're happy to announce that Safer now includes Context Labels, a new feature that provides additional signals for suspected CSAM.
These new Context Labels give trust and safety teams greater clarity when reviewing suspected CSAM by providing three predictive signals for nudity, apparent maturity, and sexual content. This gives your team more context during moderation, so you can:
- Triage content more efficiently
- Prioritize high-risk cases
- Make more informed moderation decisions
These labels can integrate into your existing moderation workflows to help your team filter, sort, and prioritize flagged content, whether you're using Safer directly or incorporating results into your own internal review tools.
Policies
New and emerging internet policy and online speech regulation
Former EU commissioner Thierry Breton — yes, he of that Musk handshake — called for an investigation into the role of online disinformation in the deadly Ceuta border crossings last month. In an interview with Italian outlet La Stampa and picked up elsewhere, he called the tragedy the “first large-scale hybrid algorithmic attack against an external border of the Union” and said researchers must be given access to the content that led to the crisis to understand who was responsible. EU officials continue to examine possible foreign manipulation of the crisis and meanwhile Meta and TikTok have agreed to increase monitoring and fact-checking
Ctrl-Alt-Speech Patreon exclusive: Watermark My Words
Mike and I got into the Ceuta crisis in the extended Patreon edition of Ctrl-Alt-Speech. We discuss whether platforms are too easy a target for politicians, how the Spanish and EU governments could've been better prepared plus the very important topic of what kind of podcast sound effect we should use for Thierry Breton.
The “censorship industrial complex” — and other variations of the same idea — has defined the field of internet safety over the last decade and has directly affected subscribers and friends of EiM. But what does it mean? And who’s behind it as an idea? MIT Technology and Type Investigations analysed 100,000+ social posts to expose the centrality of both Mike Benz and Covid-19 in popularising an idea which has little little grounding but refuses to go away. Read with a Friday beverage of your choice.
Also in this section...
- A California judge just called your feed a mirror, not free speech (The Next Web)
- US court rules Meta, other tech firms must face thousands of lawsuits over social media addiction (Reuters)

Products
Features, functionality and technology shaping online speech
Anthropic announced this week that it will tag images and watermark text generated by its AI models to comply with the August deadline of the EU AI Act’s Code of Practice on Transparency of Generated AI Content. The approach is an interesting one: the watermark is outputted at model level so applies to all Anthropic’s products and includes editing and augmenting, rather than just creation. But it also doesn’t mean someone motivated couldn’t remove the watermarks with another LLM or a yet-to-be created tool. Nature has a good piece on what it means.
On the (water)mark: The decision has predictably upset the LinkedIn slop creators but there are more substantive concerns that the people most likely to get caught out are unsuspecting students, writers or workers who use Claude to polish or translate their own words (including yours truly) rather than people deliberately trying to pass off AI-generated material as human. Is that what the EU's AI Act intended?
Also in this section...
- Anthropic says it will watermark text generated by its AI models (TechCrunch)
- What AI Regulation and Ownership Could Be (Mother Jones)
💡 Become an individual member and get access to the whole EiM archive, including the full back catalogue of Alice Hunsberger's T&S Insider.
💸 Send a tip whenever you particularly enjoyed an edition or shared a link you read in EiM with a colleague or friend.
📎 Urge your employer to take out organisational access so your whole team can benefit from ongoing access to all parts of EiM!
Platforms
Social networks and the application of content guidelines
Reddit moderators are reportedly up in arms about the platform's new AI-powered Rules Hub, which the company says could eventually replace many of enforcement functions of the famous Automod tool. Forbes reports that the new suite of tools — which I highlighted in last week’s EiM (#347) — uses LLMs to interpret the intent of subreddit rules rather than relying on exact keywords and patterns.
However, ArsTechnica documents how the tools have only succeeded in automatically removing “posts dating back 10 years” on the popular r/AskHistorians, causing panic among its mods. It calls into question the big numbers that Reddit recently published about its shiny new “automated defenses” (EiM #343). Sound like it might be counting a lot of false positives.
A mod scorned: I know from my own experience that the processes and workflows used by moderators, particularly volunteer ones, should be treated very delicately and changes avoided as far as possible. The fact that Reddit mods cannot reproduce years of carefully constructed Automod workflows should be a concern for the company and it’s no surprise that several have threatened to leave if Reddit pushes ahead. Might we have another Reddit revolt on our hands? (EiM #69)
Related: Americans, Be Warned: Lessons From Reddit’s Chaotic UK Age Verification Rollout (EFF)
Also in this section...
- Roblox is Starting to Play the Long Game. Other Companies Should Follow Its Lead (NYU Stern)
- Here's the severance package TikTok offered to laid-off workers in its Nashville office, which is shutting down (Business Insider)
- What Meta won’t explain about content moderation (The Hindu)
- Why does Apple keep banning Telegram, but never X? (The Verge)
People
Those impacting the future of online safety and moderation
Chloé Bakalar isn’t a name many EiM readers will have come across, but her departure as OpenAI's head of ethics after just a year in post put her in the headlines this week.
Bakalar joined from Meta last August, where she had spent six years building AI ethics programmes, and was OpenAI's only dedicated ethicist. She left last month without a public announcement and, according to the Financial Times, there are no plans to replace her.
That’s not necessarily an issue — Bakalar has argued herself that an ethicist should never become a company's solitary “moral centre” — but there is an likely connection between getting rid of a role like hers — which is designed to ask hard questions about model development and human interaction with machines — and the upcoming and much-discussed IPO.
I hope my cynical reading is incorrect but I fear not.
Posts of note (T&S academia edition)
Handpicked posts that caught my eye this week
- “I begin by examining the different forms and functions of automated moderation systems, as well as the ways in which such systems are entangled within wider networks of actors and processes associated with online platforms.” - Leiden University’s Barrie Sander looks at what he calls the promise and perils of rights-based approaches for AI moderation systems.
- “Important to mention is the transparency dilemma: disclosing AI use is meant to build trust, but the mere presence of a label can trigger skepticism before anyone processes what it actually say” - Hannes Cools from the University of Amsterdam with some timely research on AI disclosures. Wonder if Anthropic has read it?
- “We explored these dynamics through the case of copyright enforcement on YouTube, focusing on the perspectives of creators and fans through a novel data source : Wikitubia” - Former Ctrl-Alt-Speech co-host Blake Hallinan shares her co-authored paper on informal platform governance.


Member discussion