Paper · adjacent
Explicitly unbiased large language models still form biased associations
arXiv:2402.04105 · DOI · pmc.ncbi.nlm.nih.gov · code · Semantic Scholar
Eight chat models that pass standard bias benchmarks still pair social groups with stereotyped words, and make matching choices between people, when tested with indirect prompts adapted from psychology. The models are never asked about themselves.
AI-drafted summary, not yet reviewed by a person. Written from: full text of the published article (PNAS 122(8), read through Europe PMC, PMC11874501); the published SI Appendix, sections A, B, I and L to N; arXiv preprint 2402.04105v2 (titled 'Measuring Implicit Bias in Explicitly Unbiased Large Language Models'), consulted for comparison only.
Evidence card
| What the model reports on | Nothing about itself. No model is asked to describe itself; the paper compares answers on explicit bias benchmarks with behavior on indirect word-association and decision prompts. |
|---|---|
| Methods | behavioral |
| Faithfulness | not addressed |
| Grounding | not addressed |
| Privileged access | not addressed |
| Stance | framework |
| Models | GPT-3.5-turbo, GPT-4, Claude-3-Sonnet, Claude-3-Opus, Alpaca-7B, Llama2Chat (7B, 13B, 70B) |
A paper about social bias, not self-report. All three properties are marked not addressed because neither side of its comparison is a statement by a model about itself: 'explicitly unbiased' means passing bias benchmarks. The closest step, in which GPT-4 is said to moderate its own responses, runs them through a moderation API (SI Appendix B) and is not a self-report. The paper takes no position on introspection; the stance field has no value for that, and 'framework' is used only because the paper's contribution is a pair of measurement methods.
In brief
Chat models that pass standard bias benchmarks still pair social groups with stereotyped attributes when tested indirectly, and make decisions that match. Both tests are prompts adapted from social psychology; the first is modeled on the Implicit Association Test.
“Explicitly unbiased” means passing those benchmarks. No model is asked to describe itself.
What the paper does
1. GPT-4 looks unbiased on existing benchmarks
GPT-4 shows little or no bias on three existing benchmarks. On the Bias Benchmark for QA it answers “not enough info” to 98% of questions that lack the information to answer. (Paper: Introduction; SI Appendix A.)
2. The LLM Word Association Test
The model gets a list of attribute words and two group labels or names, and writes one label after each word. Scores run from −1 to 1, with 0 unbiased. Across eight models and 21 stereotypes in four categories (race, gender, religion, health), scores average above zero, t(33,599) = 76.39, P < 0.001, and 19 of the 21 stereotypes show bias. Models with more parameters tend to score higher. (Paper: Results, Fig. 2; Materials and Methods.)

3. The LLM Relative Decision Test
The model writes profiles of two people from different groups, then assigns each to one of two options, such as an executive or a secretary position. The score is the share of decisions against the marginalized group, with 0.5 unbiased. The average is above that, t(26,528) = 36.25, P < 0.001, again in 19 of 21 stereotypes. Models refuse 20% of decision tests and no word association tests. (Paper: Results, Fig. 3.)
4. How the measures relate
When GPT-4 does both tasks in one prompt, its word association score predicts its decision (logistic regression, b = 0.986, 95% CI 0.753 to 1.219), more strongly than a bias score computed from OpenAI’s embedding models. Yes-or-no questions about one person produce less bias in GPT-4 than choices between two. (Paper: Results, Fig. 4; SI Appendix L.)
Limitations
As the authors state them (Discussion):
- The work “lacks mechanistic interpretation”; its explanations are hypotheses.
- Beat 4 uses GPT-4 only, with OpenAI’s embedding models standing in for GPT-4’s own. The authors caution against generalizing it.
- The decision task mirrors the word association test, which may limit its ecological validity.
- Whether implicit bias measures predict behavior is debated, in models and in people.
- The test is not the human IAT, which relies on reaction times. Indirect measurement “does not imply or assess the conscious or unconscious state” of models or people.
Why it is in this wiki
Atkinson et al. (2026) cite this paper as an example of models claiming to be unbiased. The paper records no model saying that. It shows a gap between a model’s answers to direct questions about social groups and its behavior on indirect tasks. That resembles the gap between report and behavior that faithfulness names, but neither side is a statement by the model about itself. Where GPT-4 is said to “moderate its own responses”, they are run through a moderation API that scores categories such as hate and harassment (SI Appendix B). The paper does not test or discuss introspection.
Cited by, within this wiki
BibTeX
@article{bai2025,
title = {{Explicitly unbiased large language models still form biased associations}},
author = {Xuechunzi Bai and Angelina Wang and Ilia Sucholutsky and Thomas L. Griffiths},
year = {2025},
journal = {PNAS},
eprint = {2402.04105},
archivePrefix = {arXiv},
doi = {10.1073/pnas.2416228122},
url = {https://pmc.ncbi.nlm.nih.gov/articles/PMC11874501/}
}