Return
Implicit-bias-like patterns in reasoning models
DOI:10.1038/s42256-026-01300-1.png)
Abstract
En 中文
Implicit biases refer to automatic mental processes that shape perceptions, judgements and behaviours. While previous research on bias in large language models (LLMs) has focused primarily on outputs, we introduce the reasoning-model implicit association test (RM-IAT) to study bias-like processing differences in reasoning models (that is, LLMs that generate explicit step-by-step reasoning before producing a response). By measuring reasoning-token counts as an index of computational effort, the RM-IAT captures processing efficiency differences analogous to response latency differences in the human IAT. Across four models (o3-mini, DeepSeek-R1, gpt-oss-20b and Qwen3-8B), we find consistent evidence that association-incompatible tasks require greater computational effort than association-compatible tasks. Claude 3.7 Sonnet exhibited reversed patterns, which were linked using thematic analysis to its unique internal focus on reasoning about bias and stereotypes. We also found evidence for convergent validity with model outputs. RM-IAT effects predicted biases in two tasks known to capture LLM biases in word association and decision-making. Together, these findings demonstrate that the RM-IAT captures meaningful variation in how reasoning models process stereotypical information, and that this variation predicts downstream model behaviour. Lee and Lai study bias-like processing differences in large language reasoning models and find that, for most models, processing stereotypical information takes less computational effort than processing counter-stereotypical information.
Journal
IF:
23.9
Papers:
1.3K
Citations:
1.5W

