Return
Do Large Language Models Bias Human Evaluations?
DOI:10.1109/MIS.2024.3415208.png)
Abstract
En 中文
This article describes an experiment using the output from two different large language models (LLMs). I investigate whether the use of LLM's to rate intellectual ideas, biases the evaluation of those ideas by their human users. I compare the human users' evaluations when presented with different evaluations from those different LLM. I find that not only do the LLM's generate different ratings for the same materials, but those different ratings and their explanations result in statistically significant different average ratings by their human users. These results suggest that LLMs can affect issues such as using LLMs to grade student or research papers or enterprises using LLM to evaluate employees, products, software or other intellectual objects.
Keywords:
Large language models
Software
Intelligent systems
Design for experiments
Human factors
Performance evaluation
Journal
IF:
6.1
Papers:
1.6K
Citations:
4.5K

