The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume L-4/W1-2026
https://doi.org/10.5194/isprs-archives-L-4-W1-2026-95-2026
https://doi.org/10.5194/isprs-archives-L-4-W1-2026-95-2026
29 Aug 2026
 | 29 Aug 2026

Aitchison-Loss Training with Geospatial Embeddings Sharpens Compositional Land-Cover Maps

Ayato Kanno and Narumasa Tsutsumida

Keywords: Compositional land-cover classification, Mixed-pixel problem, Sentinel-2, OpenEarthMap, Foundation model embeddings, Aitchison distance

Abstract. Accurate land-cover maps are essential, but medium-resolution imagery (e.g., Sentinel-2 at 10 m) often contains mixed pixels that include multiple land-cover types. Standard “hard” classification assigns one class per pixel, hiding minority classes and reducing map usefulness. Compositional classification instead estimates the proportion of each class within a pixel, preserving sub-pixel detail, but requires outputs that are non-negative and sum to one, constraints not naturally handled by typical ML/DL losses. This study proposed and evaluated a deep-learning framework for compositional land-cover estimation at 10 m resolution. It compared two input feature sets: (1) reflectance from 10 Sentinel-2 multispectral bands (B2–B8, B8A, B11, B12) and (2) Embedding V1, a 64- dimensional representation from the AlphaEarth Foundations model that integrates multi-source, multi-temporal Earth observation signals. Ground-truth composition vectors were derived from OpenEarthMap by aggregating 0.25–0.5 m labels to 10 m pixels for eight classes. Three architectures (MLP, 2D-CNN, 3D-CNN) used Softmax outputs to enforce the constant-sum constraint, and two losses (MAE vs Aitchison distance) were tested. Embedding V1 improved estimation accuracy across all model architectures compared to Sentinel-2 spectral bands alone. While 3D-CNN achieved the best performance with Sentinel-2 input (MAE: 0.1126), MLP outperformed all other architectures when Embedding V1 was used (MAE: 0.0989). Comparison of fraction maps revealed that MAE produced spatially smoothed outputs, whereas Aitchison distance yielded sharper and more realistic compositions. The combination of Embedding V1, MLP, and Aitchison distance loss achieved the best overall performance, suggesting that foundation model embeddings combined with compositional losses can improve sub-pixel land-cover estimation.

Share