arrow
Return

Evaluating Synthetic Data Generation from User Generated Text

delete2024-11-25
delete0
delete
OA
AI
J
Jenny Chim *
J
Julia Ive
M
Maria Liakata
DOI:10.1162/coli_a_00540delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
User-generated content provides a rich resource to study social and behavioral phenomena.Although its application potential is currently limited by the paucity of expert labels and theprivacy risks inherent in personal data, synthetic data can help mitigate this bottleneck. Inthis work, we introduce an evaluation framework to facilitate research on synthetic languagedata generation for user-generated text. We define a set of aspects for assessing data quality,namely, style preservation, meaning preservation, and divergence, as a proxy for privacy. Weintroduce metrics corresponding to each aspect. Moreover, through a set of generation strategiesand representative tasks and baselines across domains, we demonstrate the relation betweenthe quality aspects of synthetic user generated content, generation strategies, metrics, anddownstream performance. To our knowledge, our work is the first unified evaluation frameworkfor user-generated text in relation to the specified aspects, offering both intrinsic and extrinsicevaluation. We envisage it will facilitate developments towards shareable, high-quality syntheticlanguage data
Keywords:
SYNCHRONY

Journal

Computational Linguistics cover
Computational Linguistics
IF:
5.3
Papers:
837
Citations:
2.7K

Organization

Q
Queen Mary University London
Scholars:
2.0W
Papers: 1.5W
Citations: 327
U
university of london
Scholars:
21.5W
Papers: 19.7W
Citations: 305