1
Return

A multi-phase reference matching algorithm for bibliometric analysis: design, implementation, and evaluation

delete2026-08-08
delete0
delete
OA
AI
M
Massimo Aria *
L
Luca D’Aniello
M
Maria Spano
DOI:10.1007/s11192-026-05763-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Bibliometric analyses rely on accurate citation counts, yet bibliographic databases routinely contain variant representations of the same cited reference, differing in journal abbreviation style, author name format, punctuation, or metadata completeness, that fragment citation links and distort standard indicators such as the h-index and journal impact metrics. We propose an unsupervised, multi-phase reference matching algorithm designed to consolidate these variants without requiring training data or external authority files beyond the ISO 4 List of Title Word Abbreviations (LTWA). The pipeline operates in seven phases: (i) format detection and string normalisation, which parses heterogeneous reference styles and standardises author names, titles, and pagination; (ii) ISO 4 journal-name normalisation, which maps both abbreviated and full journal names to a canonical short form using the LTWA; (iii) exact matching on DOI identifiers and normalised reference strings; (iv) blocking by first-author surname and publication year; (v) within-block fuzzy matching that combines Jaro-Winkler similarity with agglomerative hierarchical clustering to group near-duplicate references; (vi) post-processing metadata reconciliation, which merges complementary fields across matched records; and (vii) canonical representative selection, which elects the most informative variant as the group representative. Evaluation on a synthetic benchmark of 1 064 source articles under 17 controlled perturbation scenarios, yields precision, recall, and $$F_1$$ scores above 0.95 in 15 of 17 scenarios, with the two lowest-scoring scenarios still achieving $$F_1 \ge 0.78$$ . Validation on two real-world Scopus datasets demonstrates that the algorithm reduces unique cited-reference counts by 4.6-−11.1 %, with corresponding increases in concentration-based bibliometric indicators. The algorithm is implemented in the open-source bibliometrix R package, which has been widely adopted in the scientometric community.
Keywords:
Reference matching
Bibliographic record linkage
Reference deduplication
Bibliometrics
String similarity

Journal

Scientometrics cover
Scientometrics
IF:
3.5
Papers:
8.0K
Citations:
2.2W

Organization

D
department of economics and statistics
Scholars:
13
Papers: 6
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers