The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1335-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1335-2026
23 Jul 2026
 | 23 Jul 2026

GeoOpen3D: Geometry-guided training-free open-vocabulary 3D segmentation via visual foundation models

Shuai Zhang, Zhuoxiao Li, Jing Ou, Tengxi Wang, Zhecheng Shi, and Wufan Zhao

Keywords: 3D Point Cloud Segmentation, Open-Vocabulary Segmentation, Training-Free Framework, Vision–Language Models, Urban Scene Analysis

Abstract. Open-vocabulary 3D segmentation offers an attractive alternative to closed-set scene parsing, yet directly transferring 2D vision-language models to outdoor point clouds remains difficult because projection disrupts geometric continuity and sparse sampling weakens mask quality. This paper presents GeoOpen3D, a geometry-guided and training-free framework for open-vocabulary 3D point cloud segmentation. GeoOpen3D constructs a geometry-preserving RGB-D representation through projection, super-sampling, and depth enhancement to improve alignment between 3D structure and 2D foundation models. It then combines GroundingDINO for language-driven proposal generation with SAM for mask extraction, while introducing depth-aware regularization to favor structurally coherent regions and clearer boundaries. The selected masks are back-projected to the original point cloud through pixel-to-point correspondence, yielding point-wise semantic labels without any 3D model training. Experiments on the SensatUrban dataset show that GeoOpen3D achieves 42.1% mIoU, including 98.5% IoU for buildings and 97.3% IoU for vegetation, outperforming existing training-free open-vocabulary baselines. Additional experiments on a custom island dataset further demonstrate promising transferability to unseen categories. These results indicate that geometry-guided 2D-to-3D transfer provides an effective and scalable path towards open-vocabulary understanding of large-scale outdoor scenes.

Share