The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume L-4/W2-2026
https://doi.org/10.5194/isprs-archives-L-4-W2-2026-87-2026
https://doi.org/10.5194/isprs-archives-L-4-W2-2026-87-2026
28 Sep 2026
 | 28 Sep 2026

Towards AI-Generated 3D City Models: Bridging Research and Practice

Lilli Kaufhold and Martin Kada

Keywords: Building Reconstruction, 3D, Deep Learning, Validation, Airborne Laser Scanning (ALS), Face identification

Abstract. Deep learning methods for 3D building reconstruction from airborne laser scanning point clouds report increasingly strong geometric accuracy on benchmark datasets. However, benchmark performance does not guarantee that the resulting models are usable for city modelling, energy simulation, or flood and noise analysis, which generally require reconstructed buildings to be geometrically valid solids. We comparatively evaluate two architecturally distinct learning-based methods, BWFormer (wireframe prediction) and Point2Building (autoregressive mesh generation), on their own test datasets and on the ISPRS Vaihingen benchmark as a common independent site. We contribute a wireframe-to-solid post-processing pipeline with hard planarity constraints, assess validity with val3dity against the ISO 19107 standard at the polygon and solid level, apply the identical evaluation to the ground truth itself, and measure computational cost. Our results demonstrate that benchmark accuracy and geometric validity are effectively decoupled. BWFormer generates wireframe representations that, by design, do not constitute volumetric solids. Our post-processing pipeline produces geometrically valid solid models in 69.8% of cases. In contrast, Point2Building directly predicts mesh representations that do not require post-processing but produces valid solids in only 0.6% of cases. These shortcomings can be attributed in part to the training data, since none of the evaluated datasets provides valid solids as supervisory signals, implying that geometric validity cannot be explicitly learned from them. We therefore conclude that, for learning-based reconstruction methods to yield standard-compliant city models, data representations, training objectives, and evaluation benchmarks must be redesigned to treat validity as a primary optimisation objective.

Share