The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1381-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1381-2026
23 Jul 2026
 | 23 Jul 2026

Synthetic data generation for architectural typology documentation using diffusion models

Pedro Achanccaray Diaz, Leonhard Wesche, Markus Gerke, Sebastian Hoyer, and Klaus Thiele

Keywords: system halls, deep learning, synthetic references, aerial image recognition, diffusion models

Abstract. The identification and systematic recording of industrial buildings pose significant challenges for modern monument preservation. In particular, system halls have shaped the industrial landscape since the 19th century but often elude complete documentation because of their widespread distribution. These buildings serve as vital witnesses to technical innovations and economic transformation; however, assessing their architectural value requires a comprehensive inventory to determine the rarity or preservation state of specific building types. Deep learning (DL) approaches are commonly used for the automatic recording of these buildings in aerial photographs, where the primary obstacle is the scarcity of curated training datasets. We overcome this by employing generative AI, specifically Stable Diffusion (SD), to produce synthetic data. By fine-tuning the SD model with Low-Rank Adaptation (LoRA), we successfully replicate the appearance and textures of various hall types. To resolve the spatial incoherence and geometric inaccuracies inherent in standard text-to-image generation, we integrated ControlNet. This allows for precise structural grounding using semantic masks, where specific colors represent building types, and polygon shapes define their exact locations. The resulting model generates accurate synthetic samples that maintain both spectral authenticity and an accurate spatial layout. Their usability was assessed by training a building detection model on both the real and synthetic datasets, achieving 71.9 and 66.7 mIoU, respectively. Moreover, introducing a few real samples for validation during training increased the mIoU to 82.7. The detection results demonstrate that the synthetic dataset is a reliable source for training, yielding robust generalization.

Share