<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-L-4-W1-2026-95-2026</article-id>
<title-group>
<article-title>Aitchison-Loss Training with Geospatial Embeddings Sharpens Compositional Land-Cover Maps</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kanno</surname>
<given-names>Ayato</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Tsutsumida</surname>
<given-names>Narumasa</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>Graduate School of Science and Engineering, Saitama University, Japan</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>Space Data Frontiers Research Center, Fujitsu Research, Fujitsu Limited, Japan</addr-line>
</aff>
<pub-date pub-type="epub">
<day>29</day>
<month>08</month>
<year>2026</year>
</pub-date>
<volume>L-4/W1-2026</volume>
<fpage>95</fpage>
<lpage>102</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Ayato Kanno</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/L-4-W1-2026/95/2026/isprs-archives-L-4-W1-2026-95-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/L-4-W1-2026/95/2026/isprs-archives-L-4-W1-2026-95-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/L-4-W1-2026/95/2026/isprs-archives-L-4-W1-2026-95-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/L-4-W1-2026/95/2026/isprs-archives-L-4-W1-2026-95-2026.pdf</self-uri>
<abstract>
<p>Accurate land-cover maps are essential, but medium-resolution imagery (e.g., Sentinel-2 at 10 m) often contains mixed pixels that include multiple land-cover types. Standard &amp;ldquo;hard&amp;rdquo; classification assigns one class per pixel, hiding minority classes and reducing map usefulness. Compositional classification instead estimates the proportion of each class within a pixel, preserving sub-pixel detail, but requires outputs that are non-negative and sum to one, constraints not naturally handled by typical ML/DL losses. This study proposed and evaluated a deep-learning framework for compositional land-cover estimation at 10 m resolution. It compared two input feature sets: (1) reflectance from 10 Sentinel-2 multispectral bands (B2&amp;ndash;B8, B8A, B11, B12) and (2) Embedding V1, a 64- dimensional representation from the AlphaEarth Foundations model that integrates multi-source, multi-temporal Earth observation signals. Ground-truth composition vectors were derived from OpenEarthMap by aggregating 0.25&amp;ndash;0.5 m labels to 10 m pixels for eight classes. Three architectures (MLP, 2D-CNN, 3D-CNN) used Softmax outputs to enforce the constant-sum constraint, and two losses (MAE vs Aitchison distance) were tested. Embedding V1 improved estimation accuracy across all model architectures compared to Sentinel-2 spectral bands alone. While 3D-CNN achieved the best performance with Sentinel-2 input (MAE: 0.1126), MLP outperformed all other architectures when Embedding V1 was used (MAE: 0.0989). Comparison of fraction maps revealed that MAE produced spatially smoothed outputs, whereas Aitchison distance yielded sharper and more realistic compositions. The combination of Embedding V1, MLP, and Aitchison distance loss achieved the best overall performance, suggesting that foundation model embeddings combined with compositional losses can improve sub-pixel land-cover estimation.</p>
</abstract>
<counts><page-count count="8"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>