返回
A data- and knowledge-driven framework for developing machine learning models to predict soccer match outcomes
DOI:10.1007/s10994-024-06625-9.png)
摘要
En 中文
The 2023 Soccer Prediction Challenge invited the machine learning community to develop innovative methods to predict the outcomes of 736 future soccer matches. The Challenge included two tasks. Task 1 was to forecast the exact match score, i.e., the number of goals scored by each team. Task 2 was to predict the match outcome as probability vector over the three possible result categories: victory of the home team, draw, and victory of the away team. Here, we present a new data- and knowledge-driven framework for building machine learning models from readily available data to predict soccer match outcomes. A key component of this framework is an innovative approach to modeling interdependent time series data of competing entities. Using this framework, we developed various predictive models based on k-nearest neighbors, artificial neural networks, naive Bayes, and ordinal forests, which we applied to the two tasks of the 2023 Soccer Prediction Challenge. Among all submissions to the Challenge, our machine learning models based on k-nearest neighbors and neural networks achieved top performances. Our main insights from the Challenge are that relatively simple learning algorithms perform remarkably well compared to more complex algorithms, and that the key to successful predictions lies in how well soccer domain knowledge can be incorporated in the modeling process.
Keyword:
2023 soccer prediction challenge
k-NN
Ordinal forests
Naive Bayes
Neural networks
Outcome prediction
Soccer analytics
Super league
期刊
IF:
2.9
论文数:
2.7K
被引数:
3.4W
机构
引用论文
Learning to predict soccer results from relational data with gradient boosted trees
MACHINE LEARNING
IF2.9
Home advantage in sport - An overview of studies on the advantage of playing at home
SPORTS MEDICINE
IF9.4
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead停止解释高风险决策的黑盒机器学习模型,而改用可解释的模型

