The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B3-2026
https://doi.org/10.5194/isprs-archives-XLIX-B3-2026-3-2026
https://doi.org/10.5194/isprs-archives-XLIX-B3-2026-3-2026
30 Jul 2026
 | 30 Jul 2026

A Multimodal and Multitemporal Deep Learning Semantic Segmentation Method based on Variational Autoencoder for Multimodal Remote Sensing Image Time Series

Luca Bergamasco, Marius Guichard, Mauro Dalla Mura, and Francesca Bovolo

Keywords: Variational Autoencoder, Semantic Segmentation, Multimodal Analysis, Image time series, Remote Sensing

Abstract. Multimodal Remote Sensing (RS) methodologies have been increasingly studied in recent years due to their capacity to analyze multimodal RS data acquired from different sensors, thereby providing improved temporal resolution and extracting richer information than single-modal RS data. Deep Learning (DL) methodologies have accelerated the study of multimodal RS methods, thanks to their ability to learn features during training automatically. Many multimodal DL methods exploit this capability to learn a shared domain across modalities. However, most of them struggle to align heterogeneous modalities in a common representation. For this reason, we propose a supervised multimodal DL method that analyzes image time series acquired by different sensors to perform semantic segmentation. The proposed DL method is based on a Variational Autoencoder (VAE) that models the spatio-temporal information of the multimodal input image time series, with encoders and decoders composed of 3D convolutional layers, and learns the probability distributions for each modality. The probability distributions are combined to derive a joint distribution used for semantic segmentation. Learning the joint probabilistic distribution is achieved by combining the probabilistic parameters across modalities using a Product of Experts (PoE) approach. The feature maps derived from the obtained latent space are processed through three decoders. Two decoders aim to reconstruct the input multimodal image time series. The third decoder performs a semantic segmentation based on the inputs. Experiments conducted on the MultiSenGE and Austria datasets, which comprise Sentinel-1 and Sentinel-2 image time series acquired in France and Austria and representing heterogeneous classes, yielded promising results.

Share