arrow
Return

Data Augmentation Using Large Language Models: Methods, Challenges, and Perspectives

delete2026-01-01
delete0
PRE
AI
A
Agata Kozina *
M
Michał Pikus
J
Jarosław Wąs
DOI:10.1007/978-3-032-13869-9_5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data augmentation plays a pivotal role in improving machine learning performance, especially when labeled data are limited. With the rapid advancement of Large Language Models (LLMs), their ability to generate high quality synthetic data has attracted significant attention. In this study, we examine the use of LLMs for data augmentation, particularly the Polish large-language model Bielik. We assess its proficiency in producing diverse and contextually relevant synthetic text to enrich datasets within the financial sector, using data sourced from a leasing company. Our investigation covers a variety of enhancement strategies. These include controlled text generation through paraphrasing (with adjustable prompt parameters), the addition of noise to continuous variables (with modifiable scaling), and binary variable modifications through flipping at predetermined rates. We evaluated the impact of these techniques on overall model performance while addressing challenges such as data quality, bias mitigation, and ethical considerations in integrating LLM generated data into machine learning workflows. Ultimately, this research provides forward looking perspectives on the evolving applications and potential enhancements of LLM driven data augmentation, offering valuable insights for both practitioners and researchers.
Keywords:
Data augmentation
Large Language Models
Synthetic data
Financial sector
Machine learning

Journal

E
EMERGING CHALLENGES IN INTELLIGENT MANAGEMENT INFORMATION SYSTEMS, ECAI 2025-IMIS WORKSHOP, VOL 1
IF:
0
Papers:
27
Citations:
0

Organization

A
agh university of krakow
Scholars:
1.4K
Papers: 673
Citations: 0
W
wroclaw university of economics & business
Scholars:
543
Papers: 606
Citations: 1