The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume L-4/W3-2026
https://doi.org/10.5194/isprs-archives-L-4-W3-2026-75-2026
https://doi.org/10.5194/isprs-archives-L-4-W3-2026-75-2026
29 Sep 2026
 | 29 Sep 2026

Holistic Satellite Image Super Resolution Using Large Diffuse Generative Models

Vladimir V. Kniaz, Petr V. Moshkantsev, Victor S. Aleksandrov, Artem N. Bordodymov, Vladimir A. Knyaz, and Egor R. Smirnov

Keywords: remote sensing, satellite imagery, deep learning, image super resolution

Abstract. Satellite image super-resolution (SR) is a critical task in machine vision, supporting applications such as urban planning, agricultural monitoring, and disaster response. The objective of SR is to enhance the spatial resolution of low-resolution (LR) satellite imagery, thereby extracting finer detail. Over the past two decades, numerous SR methods have been developed, primarily leveraging deep learning to achieve significant resolution enhancements. However, prevailing approaches exhibit two principal limitations. First, many methods operate via analytic mathematical techniques or rely on learned priors from training data, without explicitly incorporating structured scene knowledge (e.g., road networks, building footprints). Second, existing methods offer minimal control over the aesthetic and structural characteristics of the output during the SR process, which is a drawback for applications like digital map updating where consistency with existing geospatial databases is essential. In scenarios where rich prior semantic information about a scene is available—such as from OpenStreetMap or other GIS sources—it is natural to harness this data to guide and improve SR reconstruction.
This paper introduces a semantic super-resolution (SSR) model designed to address these gaps by integrating semantic priors directly into the SR pipeline. The approach builds upon the Flux Kontext diffusion-based generative framework, incorporating two key modifications. Firstly, an additional textual encoder is trained to convert available semantic information—such as vector data describing roads, buildings, and water bodies—into a dense textual prior that is fed into the network. Secondly, the loss function is refined via a Low-Rank Adaptation (LoRA) adapter, emphasizing semantic consistency of major features in the generated highresolution (HR) image.
The proposed SSR model was evaluated against three state-of-the-art deep learning SR baselines: EDSR, SRGAN, and ResShift, using the high-resolution aerial imagery dataset (HRAID) comprising 236 images. Qualitative assessment reveals enhanced consistency in the reconstruction of structured features like roads and buildings, attributable to the injected textual prior. Quantitative evaluation demonstrates that our model outperforms the best-performing baseline by 5% in PSNR and by 10% in FID. An ablation study systematically removing components of the SSR model confirms the necessity of both the textual encoder and the semantic-consistency-focused LoRA adapter for achieving these gains.

Share