<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-XLIX-B4-2026-679-2026</article-id>
<title-group>
<article-title>Monocular Depth Estimation from UAV Images for 3D Documentation of Architectural Heritage: A Depth Anything V2-Based Approach</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Chiabrando</surname>
<given-names>Filiberto</given-names>
<ext-link>https://orcid.org/0000-0002-4982-5236</ext-link>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Gallitto</surname>
<given-names>Francesca</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lingua</surname>
<given-names>Andrea Maria</given-names>
<ext-link>https://orcid.org/0000-0002-5930-2711</ext-link>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Manca</surname>
<given-names>Stefania</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Martino</surname>
<given-names>Alessio</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Matrone</surname>
<given-names>Francesca</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Spadaro</surname>
<given-names>Alessandra</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>DAD, Politecnico di Torino, Viale Pier Andrea Mattioli 39, 10125 Torino, Italy</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>DIATI, Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129 Torino, Italy</addr-line>
</aff>
<pub-date pub-type="epub">
<day>04</day>
<month>08</month>
<year>2026</year>
</pub-date>
<volume>XLIX-B4-2026</volume>
<fpage>679</fpage>
<lpage>686</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Filiberto Chiabrando et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B4-2026/679/2026/isprs-archives-XLIX-B4-2026-679-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/XLIX-B4-2026/679/2026/isprs-archives-XLIX-B4-2026-679-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B4-2026/679/2026/isprs-archives-XLIX-B4-2026-679-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/XLIX-B4-2026/679/2026/isprs-archives-XLIX-B4-2026-679-2026.pdf</self-uri>
<abstract>
<p>Monocular depth estimation (MDE) has reached notable maturity in computer vision, yet its application to UAV-based architectural heritage documentation remains underexplored. This study assesses whether the depth foundation model Depth Anything V2 can be transferred from terrestrial to aerial imagery. The analysis relies on MDE4BH, a benchmark of over 3,000 UAV images covering ten heterogeneous heritage scenarios (urban areas, fa&amp;ccedil;ades, towers, villas, domes, and archaeological sites). Masked photogrammetric depth maps serve as metric reference for calibration, validation, and supervised retraining. Two baseline configurations are evaluated: a relative model with scene-specific linear rescaling and the direct application of the metric model. The rescaled relative model shows acceptable performance in several subsets, whereas the metric model exhibits systematic bias, weak consistency, and scale collapse due to domain shift between terrestrial training data and aerial acquisition geometry. To address these limitations, a two-step fine-tuning strategy is introduced, focusing on the decoder and regression head. The first stage uses mainly oblique UAV images; the second integrates oblique and nadir views to improve viewpoint generalization. The adapted model significantly reduces bias and enhances metric stability across the benchmark. However, residual errors remain spatially structured, with clustering and recurrent artefacts near object boundaries, multi-level roofs, and radiometrically heterogeneous surfaces. Although accuracy is still insufficient for demanding metric applications, the results support the use of MDE as a complementary source for thematic interpretation, scene understanding, robotics, navigation, and related tasks where strict geometric precision is not required.</p>
</abstract>
<counts><page-count count="8"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>