AI governance

Oct
02

Anti-terrorism 'accountability theatre', Meta's Muse problem and Jess Davies vs xAI

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #357
8 min read
Sep
28

[Revisited] Is it time to unite T&S and AI governance?

AI is now everywhere, but the teams managing its risks still sit apart from the people who have spent decades dealing with online harms. Eighteen months after first arguing they belong together, Alice updates her case.
7 min read
Sep
25

The other AI alignment problem, Discord takes a guess and new Ofcom probe

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #356
6 min read
Sep
18

EU teen ban lands (with a twist), AI labs' safety PR blitz and a T&S story worth telling

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #355
7 min read
Sep
08

The "rogue agent" narrative is insufficient

From a covert OpenAI message board to a hacked gym booking, this year's agentic AI incidents look like autonomy gone wrong. But that's not the full story.
5 min read
Aug
14

Watermarked words, OpenAI's very ethical exit another Reddit revolt?

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #351
7 min read
Jul
03

AI replaces moderators (again), governments replace nuance and Kim Cameron was right

The need-to-know developments shaping digital platforms, online speech and internet safety - edition #345
6 min read
Apr
21

Is this what assessing risk *actually* looks like?

Regulators have spent years trying to get platforms to anticipate harm before it happens. Anthropic’s Mythos release suggests some AI labs may already be adopting similar principles.
4 min read
Apr
01

T&S is political. Fund it like it is.

The Trust & Safety Summit reminded us that, if T&S is central to how platforms govern speech, behaviour and risk, it should be treated as a strategic function rather than a cost centre.
9 min read
Feb
27

Anthropic plays defence, Discord pleas for forgiveness and Reddit plans to appeal

The week in content moderation - edition #328
5 min read