arrow
Return

Collecting protest event data using natural language processing models

delete2025-12-01
delete0
PRE
AI
B
Bogdan Mamaev *
DOI:10.1080/1060586X.2025.2600874delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article presents the Russian Contentious Events Dataset (RCED), a comprehensive dataset of contentious events in Russia from 2010 to 2023. It addresses the challenge of collecting protest event analysis (PEA) data by using social media and natural language processing (NLP) models to automatically identify and analyze reports of such events. The article details a methodological workflow that includes event classification, named entity recognition, duplicate removal, and post-processing. Analysis of the generated dataset reveals longitudinal protest trends in Russia. Using summaries of the original tweets, the paper demonstrates how event classification and spatial clustering can be used to analyze contention at both federal and regional levels, identifying significant variations across the country. This study shows that modern machine learning and language models can automate and scale PEA, improving the data collection process while reducing resource requirements. The article contributes to the literature on the automation of protest data collection and emphasizes the need for further research into how advances in NLP can be applied in this field.
Keywords:
Contentious events analysis
protest event dataset
protests in Russia
contentious action in Russia
RCED

Journal

P
Post-Soviet Affairs
IF:
3
Papers:
44
Citations:
1.0K

Organization

D
deakin university
Scholars:
1.6K
Papers: 777
Citations: 1