城市内涝AI预报模型样本规模优化及性能响应特征

Sample size optimization and performance response characteristics of ai-based urban pluvial flood forecasting models

  • 摘要: 城市内涝AI预报模型面临样本规模选择的问题,样本过多易造成计算资源浪费,样本过少则易发生过拟合。本文将水文水动力模型与AI技术结合,选取Ridge回归、K近邻、随机森林和BP神经网络4种算法,通过识别各模型的性能突变点、饱和点与最优样本量,探究各模型预报性能随样本量增加的演变规律。结果表明:适度规模的样本即可支撑模型达到高精度收敛;保证精度的前提下(ENS≥0.85且ERMS≤0.02),4种模型的最优样本量分别为100、90、50和120,分别占总样本量的59.5%、53.6%、29.8%和71.4%;在最优样本量下,数据集构建耗时较全样本分别减少5.00、5.71、8.61和3.39 h,RF的样本优化潜力最大。研究为兼顾数据集构建成本与预测精度提供了参考。

     

    Abstract: AI-based forecasting models for urban flooding face the challenge of selecting an appropriate sample size: an excessively large sample size may result in unnecessary computational costs, whereas an insufficient sample size may lead to overfitting. In this study, hydrological and hydrodynamic modeling was integrated with AI techniques, and four algorithms, namely ridge regression, k-nearest neighbors (KNN), random forest (RF), and backpropagation (BP) neural network, were employed. By identifying the performance change points, saturation points, and optimal sample sizes of the four models, the variation in forecasting performance with increasing sample size was investigated. The results showed that moderately sized datasets were sufficient for the models to achieve high and stable forecasting accuracy. Under the prescribed accuracy criteria (ENS≥0.85 and ERMS≤0.02), the optimal sample sizes for ridge regression, KNN, RF, and BPNN were 100, 90, 50, and 120, respectively, accounting for 59.5%, 53.6%, 29.8%, and 71.4% of the full dataset. Compared with the use of the full dataset, the dataset construction time at the optimal sample sizes was reduced by 5.00, 5.71, 8.61, and 3.39 h, respectively. Among the four models, RF exhibited the greatest potential for sample-size optimization. These results provide a basis for balancing dataset construction costs and forecasting accuracy in AI-based urban flood forecasting.

     

/

返回文章
返回