The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B3-2026
https://doi.org/10.5194/isprs-archives-XLIX-B3-2026-651-2026
https://doi.org/10.5194/isprs-archives-XLIX-B3-2026-651-2026
30 Jul 2026
 | 30 Jul 2026

Assessment of Different Architectures based on 2D-UNet, 3D-UNet and UNet-ConvLSTM for Land Use Land Cover Classification Using Multi-modal and Multi-temporal Satellite Images

Fatemeh Saba, Pedro Achanccaray Diaz, and Markus Gerke

Keywords: Land cover land use classification, UNet segmentation, Remote sensing dataset

Abstract. Land Use/Land Cover (LULC) classification is essential for environmental monitoring and sustainable development. However, accurately classifying complex and heterogeneous land cover types remains challenging, particularly in capturing spatial–temporal dynamics. This study evaluates different UNet-based semantic segmentation models trained on multi-temporal and multi-modal (optical Sentinel-2 and radar Sentinel-1) images for LULC classification. We assessed the performance of two 2D-UNet approaches: processing each time step separately (2D-UNet) and stacking time steps as additional channels (2D-UNet-multitemporal). Also, 3D-UNet models are explored, including a standard version and a 3D-UNet with attention block (3D-UNet-attention). Finally, UNet variants with Convolutional Long Short-Term Memory (ConvLSTM) are evaluated as well: replacing all convolutional layers with ConvLSTM (UNet-ConvLSTM-all), adding ConvLSTM at the bottleneck (3D-UNet-ConvLSTM-bottleneck), and incorporating it at the end (3D-UNet-ConvLSTM-end). The results showed that 3D-UNet models outperformed 2D-UNets, with 4–5 % improvements in mean F1-score, particularly for vegetation-related classes. The 3D-UNet-ConvLSTM-bottleneck indicated a slight improvement over 3D-UNet-ConvLSTM-end and outperformed UNet-ConvLSTM-all, demonstrating the effectiveness of combining the 3D convolutions with ConvLSTM for spatiotemporal feature extraction. Notably, the 3D-UNet-attention model achieved the best results for LULC classification, particularly for detecting and monitoring dead tree areas, which is the most challenging class in our application. Overall, this study demonstrates the superior capability of 3D-UNet architectures in capturing temporal dynamics and enhancing classification accuracy for complex LULC tasks. Leveraging freely accessible satellite data, these findings provide valuable insights for forest managers, tree mortality assessment, and the identification of effective control measures.

Share