CMLGF-LIO: A Cross-Modal Local-Global Fusion Framework for Robust LiDAR-Inertial Odometry
Keywords: Deep Learning, LiDAR-Inertial Odometry, Cross-Modal Fusion, Local-Global Interaction, Robust Localization
Abstract. Accurate and robust localization is essential for autonomous vehicles and mobile robots operating in complex, dynamic environments. However, existing learning-based LiDAR-inertial odometry (LIO) methods typically rely on simplistic weighted fusion or purely global attention, which often fail to fully exploit cross-modal complementarity. In this paper, we propose CMLGF-LIO, a cross-modal local-global fusion framework that enhances both the accuracy and robustness of LIO. At the local level, we design a Local Split-Attention (LSA) module that injects IMU-derived motion priors into local LiDAR feature groups and adaptively allocates attention weights, suppressing redundant information while preserving discriminative local geometry for fine-grained fusion. At the global level, we introduce a Global MLP-Mixer (GMM) module that aligns LiDAR and IMU token sequences and models global cross-modal interactions using an MLP-Mixer backbone. Experiments demonstrate that CMLGF-LIO is more robust than learning-based baselines under challenging conditions, and ablation studies validate the effectiveness of the proposed local-global fusion strategy.
