Geo-Visual Fusion: An Enhanced Strategy for Drone Object Detection Based on High-Definition Map Context
Keywords: UAV Object Detection, Geo-Visual Fusion, High-Definition Maps, Contextual Reasoning, Semantic Compatibility
Abstract. With the rapid advancement of Urban Air Mobility (UAM), vision-only UAV object detectors like YOLO often suffer from "context blindness" in complex urban canyons, leading to logical fallacies or missed occluded targets. To address these limitations, this paper proposes an innovative Geo-Visual Fusion (GVF) enhancement strategy. By leveraging high-definition (HD) city maps as deterministic geo-spatial priors, we introduce a Geo-spatial Contextual Reasoning (GCR) module to post-process raw visual outputs. This framework incorporates a Semantic Compatibility Matrix (SCM) to eliminate geographically implausible false positives and a Bayesian enhancement rule to boost the confidence of occluded targets. Experimental validation in the Baibuting Community, Wuhan, demonstrates that the GVF framework significantly outperforms the baseline YOLOv11, achieving perfect recall and precision in the test sequence. Furthermore, the 2D vector-based indexing ensures high computational efficiency for edge computing deployment on platforms like the DJI Dock 3. Finally, a closed-loop "reverse empowerment" mechanism for HD map updates is discussed. This work effectively bridges probabilistic computer vision and deterministic geospatial constraints for reliable UAV perception.
