arrow
返回

Improving Predictions by Nonlinear Regression Models from Outlying Input Data

delete2023-01-01
delete5
delete
OA
AI
W
William W. Hsieh *
DOI:10.3808/jei.202300493delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
. When applying machine learning/statistical methods to the environmental sciences, nonlinear regression (NLR) models often perform only slightly better and occasionally worse than linear regression (LR). The proposed reason for this conundrum is that NLR models can give predictions much worse than LR when given input data which lie outside the domain used in model training. Continuous unbounded variables are widely used in environmental sciences, whence not uncommon for new input data to lie far outside the training domain. For six environmental datasets, inputs in the test data were classified as outliers and non-outliers based on the Mahalanobis distance from the training input data. The prediction scores (mean absolute error, Spearman correlation) showed NLR to outperform LR for the non-outliers, but often underperform LR for the outliers. An approach based on Occam's Razor (OR) was proposed, where linear extrapolation was used instead of nonlinear extrapolation for the outliers. The linear extrapolation to the outlier domain was based on the NLR model within the non-outlier domain. This NLROR approach reduced occurrences of very poor extrapolation by NLR, and it tended to outperform NLR and LR for the outliers. In conclusion, input test data should be screened for outliers. For outliers, the unreliable NLR predictions can be replaced by NLROR or LR predictions, or by issuing a no reliable prediction warning.
Keyword:
artificial intelligence
artificial neural network
extrapolation
extreme learning machine
machine learning
nonlinear
regression
outlier

期刊

J
Journal of Environmental Informatics
IF:
5.4
论文数:
530
被引数:
661

机构

U
University of British Columbia
学者数:
7.0W
论文数: 6.1W
被引数: 8.6W
引用论文

引用论文

暂无论文信息