Evaluating deep learning models for streamflow prediction under RCP 4.5 and RCP 8.5 climate scenarios
Journal of Atmospheric and Solar-Terrestrial Physics, cilt.287, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 287
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.jastp.2026.106972
- Dergi Adı: Journal of Atmospheric and Solar-Terrestrial Physics
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Artic & Antarctic Regions, Compendex, INSPEC, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO)
- Anahtar Kelimeler: Climate change, Deep learning, Explainable AI, Gated recurrent unit (GRU), Runoff generation, SHAP, Streamflow prediction, Yamula dam basin
- Erciyes Üniversitesi Adresli: Evet
Özet
Long-term management of water in semi-arid catchments relies on projections of stream flow associated with climate change. In this research, the performance of five deep learning algorithms, namely Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Bidirectional Long Short-Term Memory (BiLSTM), Temporal Convolutional Network (TCN), and Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM), was compared for stream flow projection at the monthly level in the Yamula Dam catchment, Turkey. Four climate projections from CNRM-CM5 and EC-Earth models, with RCP 4.5 and 8.5 scenarios, have been used from 2023 until 2100. The models were calibrated and validated using data from 1980 until 2013. The outcomes revealed that recurrent networks performed better than convolutional networks for all climate projections. In recurrent networks, GRU performed best with mean R2 = 0.9245 and low spread (SD = 0.0279), followed by LSTM (R2 = 0.9108) and BiLSTM (R2 = 0.9078). TCN and CNN-LSTM performed poorer, specifically during peak stream flow conditions, with respect to the measured stream flow values (underestimated by 25-40%); they were 25-40% below the measured stream flow values. Ensemble methods have yielded little improvement (approximately +0.2%) with high correlations (>0.97) between models. Flow dissection exposed up to 8% loss in R2 values above the 90th percentile stream flow values. GRU was 33% faster than LSTM with regard to training speed while maintaining superior predictive accuracy, making it the preferred architecture for operational streamflow forecasting in semi-arid catchments. To move beyond a purely performance-oriented comparison, the predictor set was rebuilt under an explicit causality constraint, audited numerically for leakage (maximum sensitivity of any predictor to any future value: 0.000), and the resulting mapping was interpreted with Kernel SHAP over the 137-month test period of each projection. The attributions are hydrologically consistent. Explanatory mass peaks in spring (1.5-1.8 times the winter total) and under wet antecedent conditions (1.4-1.9 times the normal-condition total); long-memory descriptors (6- and 12-month antecedent rainfall, third-order flow lags) dominate when the catchment is dry and in the low-flow tercile, whereas current-month rainfall and short-term rainfall variability dominate when it is wet and in the high-flow tercile. The models withhold flow in winter and release it in spring, reproducing the snow storage-and-release regime of a 989-2985 m catchment although no snow variable was supplied, and encode a wetness threshold near 30-45 mm month−1 in the 6-month antecedent rainfall above which the flow response amplifies sharply.