Comparison of Different Object Detection Methods for Automatic Facade Enrichment of Existing Building Models from Aerial Images
Keywords: Object Detection, Window, Facade Modeling, Aerial Images, Neural Network, Segment Anything
Abstract. This study investigates the enrichment of existing building models using deep learning-based window detection from oblique aerial imagery acquired by a high-end multi-camera sensor system. While many cities maintain LOD2 building models at Level of Detail 2, higher levels of detail require the integration of facade elements such as windows. Three detection strategies are evaluated using 3D reference building models to assess accuracy and completeness. The test site is located in Vienna and consists of multiple large residential buildings with varying facade characteristics. The evaluated methods include zero-shot object detection with Grounding DINO combined with Segment Anything Model 2, applied to both oblique images and facade orthophotos, as well as a SAM2-UNeXT network requiring minimal training. Results indicate that zero-shot detection on orthophotos achieves the best performance, with a precision of 0.95 and an F1 score of 0.85. In contrast, the SAM2-UNeXT approach shows lower precision and F1 scores but slightly higher recall. The investigation shows that detection performance is influenced by facade viewing angles. Steeper viewing angles generally improve detection quality but increase susceptibility to occlusions, particularly in dense urban environments. The article concludes with a detailed outlook on future work, including the extension of the approach to more complex three-dimensional building structures.
