The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-327-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-327-2026
23 Jul 2026
 | 23 Jul 2026

Unifying Street Scene Point Cloud Semantic Segmentation with Deformable Mesh-based Neural Representation

Yuzhou Zhou

Keywords: Point cloud, semantic segmentation, scene understanding, neural representation, street scene reconstruction

Abstract. Accurate semantic segmentation of urban point clouds is important for applications such as urban planning and autonomous driving. Recently, neural scene representations have been extended to merge semantic information across modalities and spatial dimensions. While 3D Gaussian Splatting (3DGS) enables efficient and high-quality reconstruction, its semantic understanding performance in street scenes is influenced by trajectory-constrained viewpoints, where Gaussian densification introduces occlusions and semantic ambiguity. This paper explores the use of NeRF-based neural representation for street scene point cloud semantic segmentation. Specifically, deformable neural mesh primitives (DNMPs) are used to compactly represent spatial geometry and simplify ray sampling. Then, neural fields including density, RGB, and semantics are constructed based on mesh vertex feature interpolation and MLPs. The sampled neural field values are accumulated via ray rendering and supervised using original images and corresponding semantic label maps generated by pre-trained models. Point cloud semantics are then predicted by interpolating neighboring samples within the learned field. The method is validated on the KITTI-360 and Waymo datasets. Results show that the proposed approach achieves improved semantic segmentation performance while maintaining competitive rendering quality, and supports both novel view synthesis and semantic rendering.

Share