The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-223-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-223-2026
23 Jul 2026
 | 23 Jul 2026

Towards Open-Vocabulary ALS Point Clouds Semantic Segmentation: An Empirical Study

Yanghong Lin, Tianyu Li, Shudong Zhou, Jingru Zhang, Li Fang, and Wei Yao

Keywords: Visual Foundation Model, ALS Point Cloud, Semantic Segmentation, Open-vocabulary, Zero-shot

Abstract. While deep learning has advanced ALS point cloud semantic segmentation and achieved impressive results, most methods rely on predefined label sets and lack ability to recognize arbitrary categories. Recently, the visual foundation models (VFMs) has garnered significant attention, due to remarkable zero-shot generalization capabilities by leveraging open-set knowledge. However, adapting these models to large-scale ALS point clouds remains largely unexplored and highly challenging. In addition, the frequent absence of well-aligned synchronously acquired images further hinders the application of 2D VFMs in ALS point clouds. To bridge these gaps, we developed a zero-shot, open-vocabulary semantic segmentation framework for ALS point clouds based on 2D-3D transfer, utilizing three types of VFMs. We employed a combination of VFMs, including source models pre-trained on natural imagery and models fine-tuned on remote sensing data, to investigate the generalization capabilities of VFMs in inherent domain gap between natural and aerial imagery. Besides, we further introduce an adaptive global view projection module that derives optimal virtual camera poses and field-of-view (FOV) from scene extents, effectively enabling the application of 2D VFMs even in the absence of original imagery. Quantitative evaluations on the Vaihingen dataset indicate that methods trained solely on natural images achieve segmentation accuracy scores of 72% (roof) and 59% (tree) for common classes but struggle with rare categories such as powerline. GSNET improves performance across most categories, highlighting importance of domain adaptation. Evaluation on the SUM dataset reveals that our approach effectively identifies large-scale urban elements (exceeding 60% precision for buildings) without high-quality, well-aligned imagery.

Share