Using textureless, low-detailed 3D city models for visual localization
Keywords: Visual localization, CityGML, 3D city models, feature matching
Abstract. Accurate camera pose estimation in urban environments remains challenging when reference imagery is generated from low-detailed, textureless 3D city models and must be matched against real world imagery. In this work we (i) extend our existing iterative object-basesd visual localization approach with an additional semantic feature and (ii) conduct a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs. As a first step to close the domain gap, we augment our iterative object-based visual localization pipeline with semantic masks derived from a pretrained semantic segmentation model. Intersection-over-Union between query and rendered masks is incorporated into the matching score, leading to a better pose accuracy. For the baseline study, we use a range of feature matching techniques: handcrafted (SIFT, AKAZE, ORB, FAST), learned detectors (XFeat, Key.Net, DeDoDe, DISK, AffNet), learned descriptors (XFeat, DISK, DeDoDe, HardNet), learned matchers (LightGlue, LoFTR), the line matcher SOLD2, and the learned matchers MINIMA-RoMa, MINIMA-LoFTR, MINIMA-XoFTR, and MatchAnything, which were trained on cross-modality datasets. The cross-modality focused matchers achieved the best results. For 20% 10%, 9%, and 7% of the evaluated query images the estimated camera pose had a translation error less than 5m and a rotation error less than 5◦. In this context, the other methods were only able to achieve a maximum success rate of 1.4%.
