1
Return

Evaluation and prediction of seasonal variation in surface water chemistry for beneficial use of drinking, irrigation, and industrial in Mahanadi River of Paradip area, Odisha, India, using an GIS: integrated water quality indices, and machine learning approaches

delete2026-08-12
delete0
delete
OA
AI
D
Dr. Abhijeet Das *
DOI:10.1007/s13201-026-02933-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Surface water is the primary source of drinking water in the Mahanadi River and its distributaries, and it is used for agricultural, industrial as well as domestic purposes. However, increasing urbanization, agricultural activities, and seawater intrusion have raised concerns about groundwater quality. This study evaluates the surface water quality (WQ) in nine selected locations by analysing twelve physicochemical parameters, computing the Weighted Arithmetic (WA) Water Quality Index (WQI), Numerow’s Pollution Index (NPI), Overall Index of Pollution (OIP), Synthetic Pollution Index (SPI), and applying Multivariate techniques namely, Corelation, Cluster Analysis (CA), and Principal Component Analysis (PCA), in order to identify key pollutants. Again, the current work is intended to create and review four machine learning (ML) models, namely, Support Vector Machine (SVM), Random Forest Model (RFM), Gradient Boosting Machine (GBM), and Extreme Gradient Boosting (XGB), for surface water quality prediction and creating spatial water quality maps to direct conservation initiatives in the heavily urbanized and polluted region being monitored. On the basis of physicochemical results, BOD, PO43−, EC, TDS, and F− were major contributors to surface water pollution, particularly in urban and agricultural areas. Findings revealed that WA WQI values ranged from 40.36 to 176.13, indicating that water quality varied from good to unsuitable category. The NPI indicated deteriorated water quality, accounting 77.78%. Considering OIP, acceptable water quality was detected around 33.33% and 66.66% of water samples renders poor class, indicating a significant risk of microbial contamination and the necessity for appropriate treatment prior to human consumption. SPI results indicate that 55.56% of the samples were classified as “good” (SPI < 0.5), could be consumed for drinking, and remaining 44.44% were classed as polluted. The surface water in the area is appropriate for use for irrigation at a variety of locations, depending on the calculated indices of SAR (5.88–35.66), MH (41.22–76.55), PI (45–85), % Na (15.22–92), and KI (0.43–2.5). For the majority of samples, the predominant water species in the region, as determined by Piper categorization, is in the category indicated as Mg2+–Cl−–SO42− water type, resulted from the dissolution of magnetite, SO42−, and Cl− minerals in the aquifer area. Gibb’s plot proven that rock-water interaction dominated the sources of surface water's chemical components. Correlation matrix describes the relationship between water quality parameters, allowing for the detection of natural or anthropogenic influences on subsurface water. Three cluster groups representing varying contamination levels were generated from nine sampling sites using twelve water quality metrics. According to their respective Eigen values, the three PCA components that were retrieved explained 56.65%, 19.90%, and 11.97% of the overall variance. The PCA results demonstrated that TDS, TH, SO42−, NO3−, and PO43− were the dominant contributors to water quality deterioration. This finding indicates that these physicochemical parameters play a critical role in influencing the overall variability and degradation of water quality in the study area.Referring to ML approaches, six statistical indicators such as Willmott's Index (WI), Nash Sutcliffe model efficiency coefficient (NSE), percent bias (PBIAS), mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R2), as well as graphical representations of radar charts and Taylor diagrams, were used to evaluate the model's performance. The combined use of statistical and graphical evaluation methods provides a comprehensive assessment of the predictive accuracy, reliability, and consistency of the developed models. During testing set, the results revealed that the performance of the RFM model (WI = 0.85, NSE = 0.94, R2 = 0.93, PBIAS = 12.02, MAE = 45.91, and RMSE = 111.43) was superior to the SVM, GBM, and XGB models for prediction of WQI. The higher predictive performance of RFM suggests its greater capability to capture the complex and nonlinear relationships among the water quality parameters influencing WQI.Interestingly, the SVM model shows significantly worse performances in predicting the WQI. This comparatively lower performance may be attributed to the model′s limited ability to effectively represent the complex interactions and variability present within the water quality dataset. This research highlights the effectiveness of machine learning in monitoring water quality and offers insights for sustainable resource management. The findings therefore demonstrate the potential of integrating machine learning with conventional water quality assessment techniques to support efficient monitoring and informed decision-making for long-term water resource conservation. It concludes with practical recommendations for targeted pollution control and the adoption of artificial intelligence (AI)—driven monitoring systems. Future studies should integrate real-time data and seasonal variations for enhanced understanding of water quality trends in the river basin.
Keywords:
Mahanadi River
Machine learning
Water quality
Multivariate
Artificial intelligence

Journal

A
Applied Water Science
IF:
5.7
Papers:
2.2K
Citations:
1.2W

Organization

D
Department of Civil Engineering
Scholars:
629
Papers: 279
Citations: 2
Cited Papers

Cited Papers

Citing Papers

Citing Papers