Return
Cross-Model Adjudication for Bias Mitigation in Large Language Models
X
C
W
Y
Y
Y
DOI:10.1109/JSTSP.2026.3662478.png)
Abstract
En 中文
With the increasing adoption and prevalence of large language models (LLMs), concerns regarding their inherent biases have become paramount. Existing approaches often rely on fixed datasets and metrics for bias detection and subsequent fine-tuning for mitigation. However, relying solely on such fixed datasets and metrics, akin to standardized exams, may be insufficient due to their inherent inflexibility (e.g., fixed testing content) and susceptibility to gaming. Furthermore, these benchmarks risk contamination if inadvertently included in training data. Drawing an analogy to human peer review and collaborative learning, this paper introduces a novel Cross-Model Adjudication Framework (CMAF) for detecting and mitigating biases in LLMs. We implement a distributed peer-review mechanism where four state-of-the-art models (Qwen2.5-7B, DeepSeek-7B-chat, Gemma2-9B, LLaMA3.1-8B) critically evaluate each other’s responses to prompts from the HolisticBias dataset. The consensus-derived, low-bias outputs are then utilized for parameter-efficient fine-tuning. Our method achieves a bias reduction of up to 12.3%, measured by the statistical significance of token likelihood differences across demographic groups, and yields another comparable B-score metric against commercial LLMs, while preserving core task performance and maintaining minimal inference latency overhead post-fine-tuning.
Keywords:
Large language models
bias mitigation
cross-model adjudication
Journal
IF:
13.7
Papers:
1.9K
Citations:
1.1W
