The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-79-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-79-2026
22 Jul 2026
 | 22 Jul 2026

ML-MIFD: Multi-Level Multimodal Invariant Feature Descriptor

Zening Wang, Haoyu Guo, Yongxiang Yao, Yongjun Zhang, Peihao Wu, and Yi Wan

Keywords: multimodal image matching, invariant feature descriptor, multi-level representation, phase congruency

Abstract. Existing cross-modal feature descriptors are typically limited to single-level local statistics, leading to insufficient structural representation and degraded matching performance under nonlinear radiometric differences and strong noise. To address this limitation, we propose a Multi-Level Multimodal Invariant Feature Descriptor (ML-MIFD). The proposed method introduces a hierarchical descriptor to jointly capture structural information at global, mid-level, and micro-level scales, thereby enhancing robustness in cross-modal representation. First, a modality-invariant feature space is constructed based on phase congruency, where stable keypoints are detected using the FAST operator. Then, a coarse-to-fine hierarchical encoding strategy is employed to integrate multi-scale structural information by cascading three complementary layers: a global contour layer, a mid-level spatial layout layer, and a micro-level texture layer. This multi-level design enables the descriptor to simultaneously preserve global structural consistency and local discriminative details. In addition, a dominant orientation assignment is introduced to ensure rotation invariance, and a feature value truncation strategy is adopted to suppress the influence of local extreme noise. Finally, initial correspondences are established using a bidirectional matching scheme, followed by outlier rejection with the Fast Sample Consensus (FSC) algorithm, resulting in improved matching accuracy and reliability. Extensive experiments on six cross-modal remote sensing scenarios demonstrate that the proposed method consistently outperforms state-of-the-art approaches in terms of matching success rate, number of correct matches, and matching accuracy, achieving up to 100% matching success rate.

Share