<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-L-4-W3-2026-75-2026</article-id>
<title-group>
<article-title>Holistic Satellite Image Super Resolution Using Large Diffuse Generative Models</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kniaz</surname>
<given-names>Vladimir V.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Moshkantsev</surname>
<given-names>Petr V.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Aleksandrov</surname>
<given-names>Victor S.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Bordodymov</surname>
<given-names>Artem N.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Knyaz</surname>
<given-names>Vladimir A.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Smirnov</surname>
<given-names>Egor R.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>State Research Institute of Aviation Systems (GosNIIAS), Moscow, Russia</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>Moscow Institute of Physics and Technology (MIPT), Moscow, Russia</addr-line>
</aff>
<pub-date pub-type="epub">
<day>29</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>L-4/W3-2026</volume>
<fpage>75</fpage>
<lpage>82</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Vladimir V. Kniaz et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/L-4-W3-2026/75/2026/isprs-archives-L-4-W3-2026-75-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/L-4-W3-2026/75/2026/isprs-archives-L-4-W3-2026-75-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/L-4-W3-2026/75/2026/isprs-archives-L-4-W3-2026-75-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/L-4-W3-2026/75/2026/isprs-archives-L-4-W3-2026-75-2026.pdf</self-uri>
<abstract>
<p>Satellite image super-resolution (SR) is a critical task in machine vision, supporting applications such as urban planning, agricultural monitoring, and disaster response. The objective of SR is to enhance the spatial resolution of low-resolution (LR) satellite imagery, thereby extracting finer detail. Over the past two decades, numerous SR methods have been developed, primarily leveraging deep learning to achieve significant resolution enhancements. However, prevailing approaches exhibit two principal limitations. First, many methods operate via analytic mathematical techniques or rely on learned priors from training data, without explicitly incorporating structured scene knowledge (e.g., road networks, building footprints). Second, existing methods offer minimal control over the aesthetic and structural characteristics of the output during the SR process, which is a drawback for applications like digital map updating where consistency with existing geospatial databases is essential. In scenarios where rich prior semantic information about a scene is available&amp;mdash;such as from OpenStreetMap or other GIS sources&amp;mdash;it is natural to harness this data to guide and improve SR reconstruction.&lt;br /&gt;This paper introduces a semantic super-resolution (&lt;code&gt;SSR&lt;/code&gt;) model designed to address these gaps by integrating semantic priors directly into the SR pipeline. The approach builds upon the &lt;code&gt;Flux Kontext&lt;/code&gt; diffusion-based generative framework, incorporating two key modifications. Firstly, an additional textual encoder is trained to convert available semantic information&amp;mdash;such as vector data describing roads, buildings, and water bodies&amp;mdash;into a dense textual prior that is fed into the network. Secondly, the loss function is refined via a Low-Rank Adaptation (LoRA) adapter, emphasizing semantic consistency of major features in the generated highresolution (HR) image.&lt;br /&gt;The proposed &lt;code&gt;SSR&lt;/code&gt; model was evaluated against three state-of-the-art deep learning SR baselines: EDSR, SRGAN, and ResShift, using the high-resolution aerial imagery dataset (&lt;em&gt;HRAID&lt;/em&gt;) comprising 236 images. Qualitative assessment reveals enhanced consistency in the reconstruction of structured features like roads and buildings, attributable to the injected textual prior. Quantitative evaluation demonstrates that our model outperforms the best-performing baseline by 5% in PSNR and by 10% in FID. An ablation study systematically removing components of the SSR model confirms the necessity of both the textual encoder and the semantic-consistency-focused LoRA adapter for achieving these gains.</p>
</abstract>
<counts><page-count count="8"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>