The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume L-4/W2-2026
https://doi.org/10.5194/isprs-archives-L-4-W2-2026-25-2026
https://doi.org/10.5194/isprs-archives-L-4-W2-2026-25-2026
28 Sep 2026
 | 28 Sep 2026

Using AI to assign building archetypes to individual buildings for enhanced resolution and accuracy in urban digital twin applications

Matthias Betz, Robert Otto, and Bastian Schröter

Keywords: Simulation, Machine Learning, Urban Digital Twin, Building Archetypes, CityGML

Abstract. Urban digital twins benefit from granular, building-level data, yet classifying city quarters into meaningful typologies remains a challenge. This paper presents a machine learning-based approach to automatically assign a building archetype to individual buildings using 3D CityGML data including building functions as the sole input. The archetypes range from detached single-family housing to industrial and business parks, and are defined by geometric properties and locational context rather than socio-economic factors, ensuring broad applicability across research domains. Because the method relies solely on widely available CityGML data, it is both portable and sector-agnostic, making it suitable for applications in energy planning, mobility research, and urban policy analysis. 
The classification pipeline is built around the urban energy simulation platform SimStadt, which extracts per-building properties from CityGML files including building height, footprint, volume, storeys, roof type, usage, and statistically derived household and occupancy data. Crucially, the approach also incorporates neighborhood context by aggregating surrounding building characteristics within a 100-meter radius. These features feed a TensorFlow-based deep learning model trained on 111 manually labeled buildings. Applied to a test dataset of 17,039 buildings from the Stuttgart region, the model completed classification in just over two minutes. Validation results show strong overall performance, with a weighted average F1 score of 82%. These results are, however, a proof of concept rather than a conclusive assessment, given the limited size of the validation dataset. Next steps are an enhanced training and validation dataset or cross-validation with other approaches, among others.

Share