The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-665-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-665-2026
23 Jul 2026
 | 23 Jul 2026

GeoRGMAE: Geospatially Guided Masked Autoencoders for Building Segmentation

Tuğba Eraslanoğlu, Guneet Mutreja, Martin Kada, and Ksenia Bittner

Keywords: Semantic Segmentation, Building Segmentation, Masked Autoencoders, Masked Image Modelling, Remote Sensing

Abstract. Accurate building segmentation from high-resolution aerial imagery is essential for various urban applications such as digital twins, geographic information system (GIS), and flood risk modelling. However, conventional supervised deep learning approaches require large amounts of pixel-level annotations, which are costly and time-consuming to obtain for large remote sensing datasets. To address this limitation, self-supervised learning (SSL) has recently emerged as an effective paradigm for learning visual representations from unlabeled data. In particular, masked autoencoders (MAE) have demonstrated strong performance by reconstructing masked image patches during pretraining. Nevertheless, conventional MAE frameworks rely on random masking strategies that ignore the spatial structure and semantic importance of regions in high-resolution remote sensing imagery. In this study, we propose GeoRGMAE, a geospatially guided masked autoencoder pretraining strategy for building segmentation. Unlike standard MAE, which rely on random masking, our approach leverages building footprint annotations available during pretraining to guide the masking process while preserving the original reconstruction objective. We introduce three masking strategies -core, balanced, and density-aware masking- that prioritize semantically relevant building regions under the varying urban densities. The core strategy focuses on building interiors, the balanced strategy distributes masking between buildings and background, and the density-aware adapts masking based on scene-level building density. Experiments on the Roof3D and WHU Building datasets demonstrate consistent, though modest, improvements over standard MAE pretraining, with the most effective masking strategy depending on dataset characteristics. These findings suggest that incorporating geospatial priors into masked image modelling (MIM) can improve representation learning for downstream building segmentation tasks.

Share