Abstract
Storm surge current-time estimations are strongly influenced by recent water-level conditions, while the temporal dependence among samples from the same typhoon event can complicate the assessment of model generalization to unseen events. This study evaluated storm surge estimation at the Lianyungang tide gauge using 2168 samples from 28 typhoon events during 2002–2024. A 41-feature extreme gradient boosting (XGBoost) model integrating historical surge and physics-motivated information was evaluated using fully nested event-grouped cross-validation, with complete outer-test typhoon events excluded from hyperparameter optimization, early stopping, and model fitting. The 41-feature XGBoost model achieved a mean absolute error (MAE) of 0.083 m, a root mean square error (RMSE) of 0.128 m, and a coefficient of determination (R2) of 0.794. On the common valid sample subset, its RMSE was 9.15% lower than that of Persistence, and it achieved a lower event-level RMSE in 18 of the 28 independently held-out typhoon events. Controlled information-source experiments showed that historical surge information accounted for most of the aggregate predictive skill, whereas adding core storm-state variables and the complete set of physics-motivated descriptors produced only limited changes in overall performance. The relative improvement over Persistence was larger for samples in which the shortest available historical surge lag was 3–6 h than for those with a 1 h lag, but this advantage did not extend consistently to the highest-surge conditions, where systematic underestimation remained evident. These results demonstrate the importance of event-independent validation for assessing machine-learning storm surge estimation and show that the model skill depends strongly on both information availability and the surge magnitude. The framework should be interpreted as a retrospective, single-station current-time estimation approach rather than an operational forecasting or extreme-surge warning system.