Return
Machine Learning Models for Bike-Sharing Demand Forecasting
DOI:10.3390/futuretransp6010026.png)
Abstract
En 中文
Bike-sharing use has been growing because it improves personal mobility, offers an alternative to walking, and strengthens connections to transit. Demand forecasting is crucial for bike-sharing services because it enables operators to anticipate empty stations and full docks, improve vehicle rebalancing and staffing, and deliver more reliable service at lower operating cost. In this paper, we propose a cluster-based, hour-ahead demand forecasting methodology that (1) groups stations into geographically coherent areas using K-means clustering method, (2) constructs hourly arrival and departure demand time series for each cluster while explicitly preserving zero-demand hours, and (3) incorporates exogenous factors such as temperature and weather-event type. We analyze multi-year trip records from Chicago's Divvy bike-sharing system (2014-2017) to characterize network expansion and assess spatial stability over time. We then use the period (1 August 2016-31 December 2017), during which the number of active stations is stable, to conduct our predictive modeling. We compare three machine learning-based predictive models-linear regression (LR), time series (TS), and random forest (RF)-and assess their out-of-sample performance using the root mean squared error (RMSE). Results show that TS and RF models consistently outperform LR, achieving up to 80% R2 values and substantially lower RMSE across all 10 clusters, with particular improvements in high-variability central areas. By forecasting net demand (arrivals minus departures) at the cluster level, the approach supports practical identification of likely surplus/deficit areas to guide rebalancing decisions.
Keywords:
demand forecasting
shared bikes
machine learning
linear regression
time series
random forest
Journal
F
IF:
1.7
Papers:
197
Citations:
386

