arrow
Return

Do Large Language Models Bias Human Evaluations?

delete2024-07-01
delete1
PRE
AI
D
Daniel E. O’Leary *
DOI:10.1109/MIS.2024.3415208delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article describes an experiment using the output from two different large language models (LLMs). I investigate whether the use of LLM's to rate intellectual ideas, biases the evaluation of those ideas by their human users. I compare the human users' evaluations when presented with different evaluations from those different LLM. I find that not only do the LLM's generate different ratings for the same materials, but those different ratings and their explanations result in statistically significant different average ratings by their human users. These results suggest that LLMs can affect issues such as using LLMs to grade student or research papers or enterprises using LLM to evaluate employees, products, software or other intellectual objects.
Keywords:
Large language models
Software
Intelligent systems
Design for experiments
Human factors
Performance evaluation

Journal

IEEE Intelligent Systems cover
IEEE Intelligent Systems
IF:
6.1
Papers:
1.6K
Citations:
4.5K

Organization

U
university of southern california
Scholars:
4.6W
Papers: 3.8W
Citations: 51