Return
MixtureMissing: An R Package for Robust and Flexible Model-Based Clustering with Incomplete Data
H
C
DOI:10.18637/jss.v115.i03.png)
Abstract
En 中文
The R package MixtureMissing performs model-based clustering on data sets with values missing at random, aiming to identify homogeneous groups of observations. In model-based clustering, the data within each cluster follow a specific distribution. In the package, 13 distributions are available, including the contaminated normal distribution, the generalized hyperbolic distribution (GHD), and 11 special or limiting cases of GHD. Notably, eight out of these 11 cases have not been formulated at the time of writing. Given a list of candidate distributions, the package can recommend the optimal distribution to employ based on a specified information criterion. In this paper, the methodological foundations and computational aspects of the package are discussed. Furthermore, important features of model fitting, model summary, and available visualization tools are thoroughly illustrated using real data sets.
Keywords:
model-based clustering
EM algorithm
outliers
skewness
missing data
contaminated normal distribution
generalized hyperbolic distribution
Journal
IF:
8.1
Papers:
616
Citations:
4.6W
