Return
A semi-supervised deep forest framework based on margin distribution optimization for tabular data
DOI:10.1016/j.ins.2026.123615.png)
Abstract
En 中文
Deep Forest (DF) is a non-differentiable deep learning model based on decision tree ensembles. As an alternative to deep neural networks, it demonstrates superior suitability for dealing with structured high-dimensional data while inherently offering interpretability. However, like many deep learning paradigms, DF often requires a substantial amount of labeled data to achieve optimal performance, posing a significant challenge in real-world scenarios where labeled samples are scarce. To address this limitation, this paper proposes a novel semi-supervised learning framework for DF, focusing on optimizing the margin distribution of both labeled and unlabeled samples. We introduce a new method to maximize the average margin of labeled data and minimize the margin variance of unlabeled data, thereby enhancing the model's generalization capability theoretically. Extensive experiments on various datasets demonstrate that our proposed semi-supervised Deep Forest (SSDF) can outperform existing semi-supervised baselines under conditions of limited labeled data.
Keywords:
Semi-supervised learning
Deep forest
Margin theory
Journal
IF:
6.8
Papers:
540
Citations:
6.2W

