Return
LLM-based feature generation from text for interpretable machine learning
DOI:10.1007/s10994-025-06867-1.png)
Abstract
En 中文
Traditional text representations like embeddings and bag-of-words hinder rule learning and other interpretable machine learning methods due to high dimensionality and poor comprehensibility. This article investigates using Large Language Models (LLMs) to extract a small number of interpretable text features. We propose two workflows: one fully automated by the LLM (feature proposal and value calculation), and another where users define features and the LLM calculates values. This LLM-based feature extraction enables interpretable rule learning, overcoming issues like spurious interpretability seen with bag-of-words. We evaluated the proposed methods on five diverse datasets (including scientometrics, banking, hate speech, and food hazard). LLM-generated features yielded predictive performance similar to the SciBERT embedding model but used far fewer, interpretable features. Most generated features were considered relevant for the corresponding prediction tasks by human users. We illustrate practical utility on a case study focused on mining recommendation action rules for the improvement of research article quality and citation impact.
Keywords:
Large language models
Feature extraction
Action rules
Journal
IF:
2.9
Papers:
2.7K
Citations:
3.4W
Organization
Cited Papers
Why was this cited? Explainable machine learning applied to COVID-19 research literature
SCIENTOMETRICS
IF3.5
A comparison of Best-Worst Scaling and Likert Scale methods on peer-to-peer accommodation attributes
no more

