Return
A multi-phase reference matching algorithm for bibliometric analysis: design, implementation, and evaluation
M
L
M
DOI:10.1007/s11192-026-05763-2.png)
Abstract
En 中文
Bibliometric analyses rely on accurate citation counts, yet bibliographic databases routinely contain variant representations of the same cited reference, differing in journal abbreviation style, author name format, punctuation, or metadata completeness, that fragment citation links and distort standard indicators such as the h-index and journal impact metrics. We propose an unsupervised, multi-phase reference matching algorithm designed to consolidate these variants without requiring training data or external authority files beyond the ISO 4 List of Title Word Abbreviations (LTWA). The pipeline operates in seven phases: (i) format detection and string normalisation, which parses heterogeneous reference styles and standardises author names, titles, and pagination; (ii) ISO 4 journal-name normalisation, which maps both abbreviated and full journal names to a canonical short form using the LTWA; (iii) exact matching on DOI identifiers and normalised reference strings; (iv) blocking by first-author surname and publication year; (v) within-block fuzzy matching that combines Jaro-Winkler similarity with agglomerative hierarchical clustering to group near-duplicate references; (vi) post-processing metadata reconciliation, which merges complementary fields across matched records; and (vii) canonical representative selection, which elects the most informative variant as the group representative. Evaluation on a synthetic benchmark of 1 064 source articles under 17 controlled perturbation scenarios, yields precision, recall, and $$F_1$$ scores above 0.95 in 15 of 17 scenarios, with the two lowest-scoring scenarios still achieving $$F_1 \ge 0.78$$ . Validation on two real-world Scopus datasets demonstrates that the algorithm reduces unique cited-reference counts by 4.6-−11.1 %, with corresponding increases in concentration-based bibliometric indicators. The algorithm is implemented in the open-source bibliometrix R package, which has been widely adopted in the scientometric community.
Keywords:
Reference matching
Bibliographic record linkage
Reference deduplication
Bibliometrics
String similarity
Journal
IF:
3.5
Papers:
8.0K
Citations:
2.2W
