Appearance-aware Scaling Diffusion Model for 3D Point Cloud Upsampling
Keywords: Airborne LiDAR, Point Cloud Upsampling, Diffusion Models, Scene Reconstruction, Deep Learning
Abstract. Airborne laser scanning (ALS) point clouds are widely used for large-scale 3D scene understanding, but acquiring dense ALS data remains costly and sparse observations often exhibit incomplete vertical structures and uneven sampling. Existing diffusion-based point cloud upsampling methods have shown promise, yet conditioning only on sparse point coordinates often leads to noisy surfaces, boundary artifacts, and limited structural recovery in under-observed regions. In this work, we propose the Appearance-aware Scaling Diffusion Model (ASDM), a conditional diffusion framework for scene-level ALS point cloud upsampling that incorporates multi-view projected depth-image cues derived directly from the input point cloud. Specifically, sparse ALS scenes are rendered from multiple virtual viewpoints to generate projected depth images, which provide complementary structural information for guiding the denoising process. These projected-image features are fused with sparse point features to improve geometric fidelity and scene-level consistency. For training and evaluation, we construct realistic sparse–dense scene pairs from aerial LiDAR data derived from the YUTO Semantic dataset using a region-disjoint split over 13 survey areas. Experiments under the ×4 upsampling setting show that ASDM outperforms recent diffusion-based baselines, achieving the best overall performance in Chamfer Distance (0.5643), JSD-3D (0.6688), F1-score (75.67), and voxelized IoU at 1m and 2m resolutions. These results demonstrate that projected-image conditioning is an effective strategy for robust airborne LiDAR scene densification.
