arrow
Return

LLM-based feature generation from text for interpretable machine learning

delete2025-10-01
delete0
PRE
AI
V
Vojtěch Balek
L
Lukáš Sýkora
V
Vilém Sklenák
T
Tomáš Kliegr *
DOI:10.1007/s10994-025-06867-1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Traditional text representations like embeddings and bag-of-words hinder rule learning and other interpretable machine learning methods due to high dimensionality and poor comprehensibility. This article investigates using Large Language Models (LLMs) to extract a small number of interpretable text features. We propose two workflows: one fully automated by the LLM (feature proposal and value calculation), and another where users define features and the LLM calculates values. This LLM-based feature extraction enables interpretable rule learning, overcoming issues like spurious interpretability seen with bag-of-words. We evaluated the proposed methods on five diverse datasets (including scientometrics, banking, hate speech, and food hazard). LLM-generated features yielded predictive performance similar to the SciBERT embedding model but used far fewer, interpretable features. Most generated features were considered relevant for the corresponding prediction tasks by human users. We illustrate practical utility on a case study focused on mining recommendation action rules for the improvement of research article quality and citation impact.
Keywords:
Large language models
Feature extraction
Action rules

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.7K
Citations:
3.4W

Organization

Cited Papers

Cited Papers

Predicting article quality scores with machine learning: The UK Research Excellence Framework
err2023-05-01
err9
errOAAI
errThelwall, Mike; Kousha, Kayvan; Wilson, Paul; Makita, Meiko; Abdoli, Mahshid; Stuart, Emma; Levitt, Jonathan; Knoth, Petr; Cancellieri, Matteo
errShare
errSave
A survey on dataset quality in machine learning
err2023-10-01
err69
errOAAI
errGong, Youdi; Liu, Guangzhen; Xue, Yunzhi; Li, Rui; Meng, Lingzhong
errShare
errSave
Explainable and interpretable machine learning and data mining
err2024-07-30
err0
errOAAI
errMartin Atzmueller; Johannes Fürnkranz; Tomáš Kliegr; Ute Schmid
errShare
errSave
Why was this cited? Explainable machine learning applied to COVID-19 research literature
err2022-04-09
err12
errOAAI
errBeranova, Lucie; Joachimiak, Marcin P.; Kliegr, Tomas; Rabby, Gollam; Sklenak, Vilem
errShare
errSave
Factors affecting number of citations: a comprehensive review of the literature
err2016-02-15
err470
PREAI
errTahamtan, Iman; Afshar, Askar Safipour; Ahamdzadeh, Khadijeh
errShare
errSave
errShare
errSave
Combining full-text analysis and bibliometric indicators.: A pilot study
err2005-03-01
err43
errOAAI
errGlenisson, P; Glänzel, W; Persson, O
errShare
errSave
no more