<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-XLIX-B2-2026-315-2026</article-id>
<title-group>
<article-title>Appearance-aware Scaling Diffusion Model for 3D Point Cloud Upsampling</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yoo</surname>
<given-names>Sunghwan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sohn</surname>
<given-names>Gunho</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>Dept. of Earth and Space Science and Engineering, York University, Toronto, ON, M3J 1P3, Canada</addr-line>
</aff>
<pub-date pub-type="epub">
<day>23</day>
<month>07</month>
<year>2026</year>
</pub-date>
<volume>XLIX-B2-2026</volume>
<fpage>315</fpage>
<lpage>320</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Sunghwan Yoo</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/315/2026/isprs-archives-XLIX-B2-2026-315-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/315/2026/isprs-archives-XLIX-B2-2026-315-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/315/2026/isprs-archives-XLIX-B2-2026-315-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/315/2026/isprs-archives-XLIX-B2-2026-315-2026.pdf</self-uri>
<abstract>
<p>Airborne laser scanning (ALS) point clouds are widely used for large-scale 3D scene understanding, but acquiring dense ALS data remains costly and sparse observations often exhibit incomplete vertical structures and uneven sampling. Existing diffusion-based point cloud upsampling methods have shown promise, yet conditioning only on sparse point coordinates often leads to noisy surfaces, boundary artifacts, and limited structural recovery in under-observed regions. In this work, we propose the &lt;em&gt;Appearance-aware Scaling Diffusion Model (ASDM)&lt;/em&gt;, a conditional diffusion framework for scene-level ALS point cloud upsampling that incorporates multi-view projected depth-image cues derived directly from the input point cloud. Specifically, sparse ALS scenes are rendered from multiple virtual viewpoints to generate projected depth images, which provide complementary structural information for guiding the denoising process. These projected-image features are fused with sparse point features to improve geometric fidelity and scene-level consistency. For training and evaluation, we construct realistic sparse&amp;ndash;dense scene pairs from aerial LiDAR data derived from the YUTO Semantic dataset using a region-disjoint split over 13 survey areas. Experiments under the &amp;times;4 upsampling setting show that ASDM outperforms recent diffusion-based baselines, achieving the best overall performance in Chamfer Distance (0.5643), JSD-3D (0.6688), F1-score (75.67), and voxelized IoU at 1m and 2m resolutions. These results demonstrate that projected-image conditioning is an effective strategy for robust airborne LiDAR scene densification.</p>
</abstract>
<counts><page-count count="6"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>