返回
A self-training algorithm based on the two-stage data editing method with mass-based
DOI:10.1016/j.neunet.2023.09.046.png)
摘要
En 中文
A self-training algorithm is a classical semi-supervised learning algorithm that uses a small number of labeled samples and a large number of unlabeled samples to train a classifier. However, the existing self training algorithms consider only the geometric distance between data while ignoring the data distribution when calculating the similarity between samples. In addition, misclassified samples can severely affect the performance of a self-training algorithm. To address the above two problems, this paper proposes a self training algorithm based on data editing with mass-based dissimilarity (STDEMB). First, the mass matrix with the mass-based dissimilarity is obtained, and then the mass-based local density of each sample is determined based on its k nearest neighbors. Inspired by density peak clustering (DPC), this study designs a prototype tree based on the prototype concept. In addition, an efficient two-stage data editing algorithm is developed to edit misclassified samples and efficiently select high-confidence samples during the self-training process. The proposed STDEMB algorithm is verified by experiments using accuracy and F-score as evaluation metrics. The experimental results on 18 benchmark datasets demonstrate the effectiveness of the proposed STDEMB algorithm.
Keyword:
Self-training algorithm
Mass-based dissimilarity
Data editing
Relative node set
期刊
IF:
6.3
论文数:
8.2K
被引数:
3.0W
机构
引用论文
How Does Nitrogen and Perenniality Influence Belowground Biomass and Nitrogen Use Efficiency in Small Grain Cereals?
Crop Science
IF0
Semi-supervised multi-label image classification based on nearest neighbor editing基于最近邻编辑的半监督多标签图像分类
NEUROCOMPUTING
IF6.5

