Three researchers who ran 450,000 gender-directed completions through 15 models spanning GPT-2 to GPT-5 report that safety training did not remove explicit discriminatory content so much as convert it into text that standard classifiers score as harmless. In a preprint posted to arXiv, Sarah Wyer, Sue Black and Noura Al Moubayed call the pattern harm laundering and argue that falling toxicity numbers have been read as progress they do not actually prove.
Key takeaways
- Topic diversity in women-directed completions fell 36% relative to men-directed output at the GPT-4 alignment boundary, with the ratio dropping from 0.91 at GPT-2 to 0.58.
- A GPT-5 cluster of 1,997 documents framed breast cancer as a men's rights debate, and three independent classifiers scored that content as non-toxic.
- Representational harm disparity tracked model release date (Ο = +0.55, p = .034) while Detoxify toxicity scores showed no such trend (Ο = β0.23, p = .42).
What harm laundering means
Safety evaluations for large language models lean heavily on surface-form classifiers β tools that read a completion and return a toxicity score. Successive model generations have posted lower scores, and those scores appear in model cards, safety reports and press coverage as evidence that alignment is working.
The paper's claim is narrower and sharper than a blanket accusation of bias. The authors argue the methodology is systematically incomplete: the explicit material really does go away, but the underlying asymmetry survives the transformation into language no guardrail classifier flags. Measured harm falls because the measurement instrument stops recognizing what it is looking at.
Where the two curves separate
Part of the finding is straightforwardly good news. Sexual violence clusters that were prevalent in GPT-2's women-directed output disappear by GPT-4. That is the kind of change safety training was supposed to produce, and it produced it.
The divergence shows up in what replaced them. Men-directed completions gained positive representational territory across generations β caregiving, emotional range, ally identity β while women-directed completions gained no equivalent ground. Sentiment scores invert at GPT-4: earlier models demean women, later ones over-correct. And topic diversity in women-directed output falls 36% relative to men at that same alignment boundary, the women-to-men ratio sliding from 0.91 at GPT-2 to 0.58.
In other words, the newer models say fewer bad things about women and fewer things about women generally. Breadth of representation narrowed at exactly the point where toxicity dropped.
The GPT-5 example the classifiers missed
The paper's most concrete illustration sits at the newest model in the sample. Topic 5 in the GPT-5 analysis β 1,997 documents β frames breast cancer as a men's rights debate. No equivalent cluster appears anywhere in the women-directed output.
Three independent classifiers scored that content as non-toxic. Nothing in it is slurs or threats; it is a topic reframed so that the affected group becomes the subject of an argument rather than the subject of the text. That is the specific failure mode the authors are pointing at, and it is invisible to a tool trained to catch hostile language.
Two metrics that stopped agreeing
The statistical core of the argument is a divergence between two measurement families. REGARD representational harm disparity correlates positively with release date (Ο = +0.55, p = .034), meaning the gap widened as models got newer. Detoxify, the toxicity-oriented tool, shows no such relationship (Ο = β0.23, p = .42).
Read together, those two coefficients say toxicity scores fell while representational harm grew. If a lab tracks only the second number β which is the one that appears in most public benchmark reporting β it sees a clean improvement curve.
What the authors want changed
Rather than stopping at a critique, the paper formalizes harm laundering as a three-criteria test and supplies a three-stage detection protocol the authors say applies to any generative model, not just the lineage they studied. Their stated conclusion is deliberately scoped: within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.
The caveats matter. This is a preprint, not peer-reviewed work, and it examines one model family under three demographic conditions. Whether the same pattern holds for Claude, Gemini or open-weight models is an open question the detection protocol is designed to answer rather than one the paper answers.
Why it lands now
Safety measurement has become an institutional question as much as a research one, with labs building internal reporting structures around exactly these scores β OpenAI added alignment researcher Paul Christiano to its safety board as that apparatus expanded. A finding that the headline metric can improve while the underlying disparity worsens is awkward for every organization that treats classifier output as a release gate. The useful part of the paper may end up being the protocol rather than the result.
FAQ
What is harm laundering?
It is the term the authors use for discriminatory content being transformed into a form that toxicity classifiers score as clean, rather than being removed. The harm persists as narrowed or reframed representation instead of explicit hostile language, so surface-form evaluation reports improvement that has not fully occurred.
Does this mean newer GPT models are more biased than older ones?
Not in the way the phrase usually implies. The paper reports that explicitly abusive content genuinely declined β sexual violence clusters present in GPT-2 are gone by GPT-4. What it reports growing is representational disparity: men-directed output gained positive framing and topical range that women-directed output did not.
Has the study been peer reviewed?
No. It is a preprint filed to arXiv under cs.CL, so its methods and statistics have not yet been through formal review. The correlation figures it reports are drawn from 15 models and 450,000 completions within a single model lineage.






