arrow
Return

Sequence analysis for large databases

delete2026-04-01
delete1
PRE
AI
S
Studer, Matthias *
S
Sadeghi, Rojin
T
Tochon, Louis
DOI:10.1332/17579597y2026d000000078delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article develops and reviews methods for the creation of sequence analysis typologies in large databases. The standard way to create such a typology relies on the computation of distances between all observations, which quickly becomes intractable with large databases, even with modern computers. We start by discussing the CLARA algorithm before extending it with methods recently proposed for sequence analysis, including fuzzy and representativeness clustering. The strengths of the approaches are assessed using simulations, which further allows drawing some practical guidelines. Next, we discuss three approaches to measure the quality of the clustering without computing all distances. The first is based on bootstrapping, while the second is based on representative sequences (that is, medoids). We then introduce a third innovative approach based on clustering stability, which further allows assessing the convergence of the clustering algorithm. The methods are illustrated through a study of family trajectories in India with more than 180,000 cases. All the methods are made available in the WeightedCluster R package.
Keywords:
sequence analysis
cluster analysis
large databases
dissimilarity
cluster validation
life courses

Journal

L
Longitudinal and Life Course Studies
IF:
0
Papers:
23
Citations:
0

Organization

U
university of geneva
Scholars:
3.6W
Papers: 2.9W
Citations: 35