arrow
Return

Subsampling-Based Consensus Hierarchical Clustering for Robust Customer Segmentation with Mixed-Type Data

delete2026-04-13
delete0
PRE
AI
M
Marefat, Nooshin
P
Purificación Galindo‐Villardón *
V
Vicente-Galindo, Purificacion
DOI:10.3390/math14081294delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Hierarchical clustering is an unsupervised framework that organizes observations according to pairwise similarity relationships. In this study, an agglomerative hierarchical approach combined with Gower dissimilarity is employed to accommodate mixed-type customer data. To address data quality issues such as missing values and outliers, Multiple Imputation by Chained Equations (MICE) and Winsorization are incorporated into the preprocessing pipeline. To validate cluster stability and identify the optimal number of clusters, we employ silhouette analysis, the Davies-Bouldin Index (DBI), the Proportion of Ambiguous Clustering (PAC), and a subsampling-based consensus clustering framework. A consensus-based hierarchical tree derived from the consensus matrix is employed to assess the robustness of the segmentation structure. The resulting clusters are further evaluated through comparisons with baseline algorithms for mixed-type data, including Partitioning Around Medoids (PAM) based on Gower dissimilarity and the K-prototypes method, together with statistical tests confirming significant behavioral differences between the identified segments. From an application standpoint, these results provide a data-driven basis for customer targeting by identifying distinct behavioral patterns, thereby supporting more effective engagement strategies and optimized resource allocation.
Keywords:
subsampling-based consensus clustering
customer segmentation
mixed-type data
hierarchical clustering
gower distance

Journal

Mathematics cover
Mathematics
IF:
2.2
Papers:
2.9K
Citations:
3.6W

Organization

U
University of Salamanca
Scholars:
1.1W
Papers: 8.0K
Citations: 9