1
Return

Comparing a Large Language Model to Human-Generated Retention Messages for a Family Healthy Weight Program

delete2026-07-29
delete0
PRE
AI
K
Kayla Norton
R
Ryan D. Burns
G
Guilherme Del Fiol
J
Jennie L. Hill
K
Kate Heelan
C
Caitlin Golden
A
Ali Mercado
P
Paul A. Estabrooks
DOI:10.1177/08901171261472636delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
<jats:sec> <jats:title>Background</jats:title> <jats:p>Text messaging can improve attendance and retention in community health programs; however, message development can be resource intensive.</jats:p> </jats:sec> <jats:sec> <jats:title>Purpose</jats:title> <jats:p>To compare human and large language model (LLM)-generated retention messages in terms of creation time, clarity, appropriateness, and alignment with behavior change principles.</jats:p> </jats:sec> <jats:sec> <jats:title>Design</jats:title> <jats:p>Mixed methods using expert message ratings and qualitative feedback from message developers.</jats:p> </jats:sec> <jats:sec> <jats:title>Setting</jats:title> <jats:p>Building Healthy Families, a family healthy weight program.</jats:p> </jats:sec> <jats:sec> <jats:title>Sample</jats:title> <jats:p>Experts in behavioral science or related fields (n = 21) and message developers (n = 3).</jats:p> </jats:sec> <jats:sec> <jats:title>Measures</jats:title> <jats:p>Matched message pairs were rated on a 5-point Likert scale for clarity, appropriateness, and alignment with Social Cognitive Theory principles.</jats:p> </jats:sec> <jats:sec> <jats:title>Analysis</jats:title> <jats:p>Wilcoxon signed-rank tests compared message ratings. Equivalence testing using a ±10% equivalence interval and 90% confidence intervals assessed similarity between message types.</jats:p> </jats:sec> <jats:sec> <jats:title>Results</jats:title> <jats:p> LLM-generated messages yielded significantly higher ratings than human-generated messages for 4 of 5 message pairs on clarity and 3 of 5 rated on appropriateness (all <jats:italic toggle="yes">P</jats:italic> ’s &lt; 0.05). Messages were statistically equivalent for 3 of 5 message pairs rated for behavior change theory alignment (all <jats:italic toggle="yes">P</jats:italic> ’s &lt; 0.05). Time to develop the human- and LLM-generated messages was approximately the same with similar averaged Flesch-Kincaid grade level scores. However, the LLM-generated messages were shorter and more concise. </jats:p> </jats:sec> <jats:sec> <jats:title>Conclusion</jats:title> <jats:p>This study suggests the potential for LLMs to create messages more efficiently and reduce, but not eliminate, workload for health promotion professionals.</jats:p> </jats:sec>

Journal

A
AMERICAN JOURNAL OF HEALTH PROMOTION
IF:
2.4
Papers:
190
Citations:
0

Organization

U
University of Utah
Scholars:
2.9W
Papers: 2.2W
Citations: 4.6W
U
University of Nebraska at Kearney
Scholars:
35
Papers: 12
Citations: 493
Cited Papers

Cited Papers

Citing Papers

Citing Papers