The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-421-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-421-2026
23 Jul 2026
 | 23 Jul 2026

VISTA-GS: MVS-Guided Virtual View Augmentation for Sparse-View 3D Gaussian Splatting

Hongsheng Huang, Yaxin Li, Shengjun Tang, Siqi Du, Mostafa Mahmoud, Mahmoud Adham, and Wu Chen

Keywords: D Gaussian Splatting, Sparse-View Novel View Synthesis, Appearance Overfitting, Virtual View Augmentation, Outof-Distribution Generalization

Abstract. 3D Gaussian Splatting (3DGS) has emerged as a leading technique for novel view synthesis (NVS), yet its performance degrades drastically under sparse-view conditions. While existing methods have sought to address this by incorporating accurate 3D geometry via Multi-View Stereo (MVS) or LiDAR priors, the view-dependent appearance parameters (i.e., spherical harmonics) remain exclusively optimized on the limited training views, leading to severe appearance overfitting. This is the fundamental reason why these geometry-enhanced methods still fail to generalize to out-of-distribution (OOD) viewpoints with large baselines, such as lane-changing trajectories in autonomous driving. To address this limitation, we propose VISTA-GS (Virtual Image Synthesis and Training Augmentation), a framework that synergizes MVS-based dense initialization with a physically-grounded virtual view augmentation strategy. Specifically, we position virtual cameras at strategic offsets around the original viewpoints and render virtual training images with binary validity masks via alpha-blending. By computing photometric losses exclusively within valid mask regions, VISTA-GS injects explicit angular constraints into the optimization process, effectively regularizing view-dependent appearance without relying on any external generative model. Experiments on the LLFF benchmark and a real-world LiDAR-scanned dataset demonstrate that our method achieves state-of-the-art NVS quality under sparse-view settings, with particularly significant improvements on challenging OOD viewpoints.

Share