Learning with Spaceborne LiDAR for Enhancement of Bare-Earth Digital Elevation Models from Global Data
Keywords: Digital elevation model, Spaceborne LiDAR, Multi-modal fusion, Depth completion, Open geospatial science
Abstract. Bare-earth Digital Terrain Models (DTMs) are fundamental to geospatial applications, yet publicly accessible global elevation products are Digital Surface Models (DSMs) with vertical accuracies of 4–15 m RMSE at 30 m or coarser resolution. Airborne LiDAR achieves sub-metre accuracy but remains costly and spatially fragmented. ICESat-2 ATL08 spaceborne LiDAR photons offer a compelling complement: unlike optical sensors, they physically penetrate dense vegetation canopies to measure ground elevation directly, and ICESat-2 acquires new measurements daily on a 91-day repeat cycle. Incorporating ATL08 into a learning framework, however, poses two challenges: extreme measurement sparsity after quality filtering and an unstructured point geometry incompatible with the dense raster structure of imagery and DSMs. We present a deep neural network that addresses both challenges by fusing remote sensing imagery, Copernicus GLO-30 DSMs, and ATL08 photons to generate 3 m bare-earth elevation from global data. A multi-scale deformable cross-attention mechanism fuses each photon with surrounding image and DSM features in their native geometry—without rasterisation—thereby turning sparse measurements into effective guidance features. A Spatial Propagation Network (SPN) then densifies predictions guided by image structure, re-anchored at photon locations at every iteration. Evaluated on the DFC30 benchmark augmented with ATL08 measurements, our method achieves consistent improvements over optical-only baselines across all vegetation classes, improving overall RMSE by 74.5% over Copernicus GLO-30 (cubically interpolated to 3 m spatial resolution), with the largest gains in dense forest canopy (83.74% RMSE reduction) where ground elevation is most difficult to recover.
