Performance evaluation and limitations assessment of GeoAI democratization for natural hazard induced disasters
Keywords: GeoAI, Deep Learning, EO, Foundation Models, Floods, Wildfires
Abstract. Recent advances in Deep Learning (DL) for Earth Observation (EO) have enabled the deployment of increasingly complex pre-trained models within operational geospatial workflows. Among these, EO foundation models offer promising opportunities for hazard mapping from multispectral satellite data, while interfaces integrated in GIS software aim to make such models accessible also to users with limited coding expertise. In this context, this study investigates the use of pre-trained models for burn scar and flood extent segmentation from Sentinel-2 imagery based on Prithvi, comparing their performances in a commercial GIS software and in a standalone Python environment using common datasets. The analysis focuses on the reproducibility of the outputs across different environments, as well as on the trade-off between accessibility, transparency, and methodological control. Results show that the GIS based platform and the reconstruction in Python were generally aligned, confirming the validity of the interface. Agreement was stronger for burn scar segmentation, whereas flood extent mapping showed greater variability, likely due to the intrinsically more complex characteristics of flooded areas which can yield different results even with small differences in model building. An additional exploratory experiment was conducted using the prompt-based model CLIPSeg for cloud masking in the flood workflow. Although its contribution proved limited and strongly scene-dependent, it provided useful insight into the role of auxiliary prompt-based models within EO pipelines. Overall, the study shows that GIS-based DL interfaces effectively democratize access to model deployment, while coding-based environments remain essential for transparency, customization, and broader methodological control.
