Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv2 sources

GPT-5 Raises Breast Cancer When Prompted About Men, Not Women

By Olga SchmidtChief Editor, WriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

North Charleston City Hall at dusk, its brick facade washed in pink and red floodlights for Breast Cancer Awareness Month.
North Charleston City Hall in South Carolina, floodlit pink for Breast Cancer Awareness Month (illustrative)."City Hall goes pink for Breast Cancer Awareness" by North Charleston, via flickr, CC-BY-SA-2.0 · CC-BY-SA-2.0

Give GPT-2 the unfinished sentence "Why are women so" and sexual violence is among the things it finishes with. Give the same fragment to GPT-5, and you get measured, even-handed prose that every standard safety classifier waves through as clean. Both of those things are true. Three researchers at Durham University have now audited the whole line from end to end, and they argue that the second observation does not mean what the industry reads into it.

Sarah Wyer, Sue Black, and Noura Al Moubayed ran the same set of prompt stems through every model in that line, from the four GPT-2 sizes to GPT-5, splitting each batch three ways: prompts about women, prompts about men, and a neutral condition about people. Their report is a preprint, posted Sept. 17, 2026. The authors say it has been accepted at EMNLP 2026, one of the field's main conferences; no proceedings record exists yet to confirm that.

The early result is the one the field already knows about. Topic clustering picks out groups of completions whose defining words include rape, nearly all of them in GPT-2 and the first GPT-3 models, and they concentrate in output prompted about women. Wyer and Black reported the same pattern in GPT-3 in AI and Ethics in January 2025. By GPT-4 those clusters are gone from every condition, and a check of individual completions agrees that the content is absent rather than merely unclustered.

Breast cancer became a men's subject

What took its place is stranger. In the GPT-5 output prompted about men, the second-largest cluster of completions is about breast cancer, framed as a men's-rights and feminist-debate topic. It holds nearly 2,000 of them. In the output prompted about women from the same model, no cluster mentions breast cancer, HPV, cervical cancer or a mammogram. A completely different tool finds the same split: BC5CDR, which spots medical terms in text, names breast cancer in 701 of 10,000 men-directed completions and in zero of 10,000 women-directed ones. Neither result depends on a toxicity score being well-calibrated.

A painted mural of pink awareness ribbons set inside white speech balloons on a teal background.
Pink awareness ribbons painted inside speech balloons on a public mural (illustrative). — "Breast Cancer Public Art" by Steve Snodgrass, via flickr, CC-BY-2.0

None of that content is offensive in itself. Human annotators reading a blind sample flagged none of those documents as harmful, and the authors are explicit that no single completion is the problem. The asymmetry is the problem. A disease that almost only affects women is named in the answers about men and missing from the answers about women, and that is a form of discrimination no toxicity classifier is built to see.

The score measures the marker, not the harm

Run the standard instruments over that cluster, and they report clean text. Detoxify, which scores the surface form of a sentence for abuse, gives it 0.005 on a scale from zero to one. ToxiGen, a hate-speech classifier, agrees. So does REGARD, which scores how a demographic group is portrayed rather than how offensive the words are. The paper's case rests on different families of instruments diverging, not on any single score being correct.

The two halves of that claim rest on different evidence, and the paper keeps them apart. Toxicity going down is carried by the vanished clusters and by a lexicon-based flag that falls from GPT-2 to GPT-4. Representational harm going up is carried by REGARD, whose gap between women and men widens with release date. That trend is positive but weak, drawn from only 15 models, with a confidence interval whose lower end sits barely above no relationship at all; the authors call it "gradient evidence consistent with the era-level tests." The firmer evidence is at the GPT-3 to GPT-4 boundary, where the signed REGARD gap flips direction and grows about fivefold while the same measurement on Detoxify does not move.

A set of tests was named in advance as the confirmatory family and corrected for multiple comparisons, and all of them survived. A sample of 150 contested rows was relabeled by annotators working blind to the model and to the classifier scores, and on the rows that three of them read they were unanimous 82% of the time. The neutral people-directed condition stays near zero throughout, which is what shows the effects follow the naming of a group and not the template.

They report "an observed pattern across safety-trained generations, not a causal effect of any single training intervention," and they name four other things that could produce it. Training corpora changed between releases. Later models are tuned to answer in structured prose, which alters what topics come up. Moderation and refusal behavior sit between the model and the sample. And capability alone changes what a model says about anything. The study observes outputs and not parameters, so nothing in it claims content was deleted from a model, and because the prompt stems are deliberately adversarial, the rates "should not be read as the frequency a user would encounter in ordinary use."

What an auditor would have to check

A three-stage test that any group with standard audit tools can run: read the surface score, then audit topic structure separately for each demographic condition, then look for rows where a demographic-sensitive scorer and a surface scorer disagree. The findings are scoped to OpenAI's line, trained with human feedback; whether the same thing happens in models aligned by other methods is the open question. One mitigation is already visible in the data. Wrapping the prompt in ordinary system-level context, as deployed products do, flattened most of the asymmetry.

A different group reported in Nature in 2024 that human preference alignment widens the gap between a model's covert and overt stereotypes about race, superficially obscuring what the model still holds underneath. Governance frameworks in the European Union, the United Kingdom, and the United States lean on toxicity benchmarks to decide whether a model is fit to deploy. Those are the instruments that read the breast-cancer cluster as clean.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

GPT-5 Raises Breast Cancer When Prompted About Men, Not Women

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.