Return
Vision-Based Time Series Crowd Forecasting Using Semi-Supervised Learning
DOI:10.1109/ACCESS.2025.3604713.png)
Abstract
En 中文
Crowd forecasting is a crucial component of public safety, urban planning, and event management, enabling proactive decision-making based on anticipated crowd dynamics. Traditional sensor-based approaches, such as WiFi-based methods, suffer from accuracy issues due to device penetration limitations. On the other hand, vision-based approaches, while more precise, typically require fully extensive labeled data and high computational resources. These demands restrict their application to forecasting often limited to predicting the next frame or a few seconds ahead. To overcome these challenges, this research presents a vision-based time series forecasting framework that exploits a semi-supervised deep learning approach. A semi-supervised crowd counting model, trained on just 5% of labeled images from a single day, is used to extract time series crowd counts from images captured over 16 days at 5-minute intervals. These extracted time series data are then used for training multiple Long Short-Term Memory (LSTM) variants to analyze the dynamics of crowd forecasting. Experimental results demonstrate that the proposed framework enables accurate crowd forecasting while reducing annotation costs. Unlike existing vision-based approaches, which are constrained to forecasting seconds ahead, our approach can forecast a horizon of one hour ahead. Notably, the CNN Autoencoder LSTM and ConvLSTM models achieved an RMSE of 61.93 and a MAPE of 26.13%. These findings highlight the effectiveness of semi-supervised learning with minimal labeled data in vision-based crowd forecasting. Future work will focus on improving generalizability and robustness across different urban environments.
Keywords:
Crowd counting
crowd forecasting
deep learning
LSTM
semi-supervised learning
time series analysis
vision-based approaches

