Visual Geometry Router
Projects VGGT-Ω features into requested cameras with confidence-aware, two-layer visibility routing.
1 XPeng Motors2 The Chinese University of Hong Kong3 Tsinghua University
Effective 21K checkpoint · Full resolution · 80 continuous frames
From six input images, VGGT-Diff generates a continuous 80-frame novel-view sequence at full resolution. The model was initialized from a 20K half-resolution checkpoint, then received only 1K steps of full-resolution continuation training.
This effective 21K checkpoint is still training. Fine details and cross-view consistency are expected to continue improving. Fully trained full-resolution checkpoints will be released on the VGGT-Diff GitHub project page.






We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing NVS methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model.
A confidence-aware Visual Geometry Router transforms 3D points and appearance-bearing tokens into query-aligned conditions while preserving front and back surface evidence. Joint target-view denoising is regularized by Point-Track Residual Consistency, which aligns predicted-clean residuals along reliable 3D tracks. Uncertainty-aware training and inference further improve robustness to imperfect geometry.
Projects VGGT-Ω features into requested cameras with confidence-aware, two-layer visibility routing.
Processes source and query slots together so requested views exchange information during denoising.
Couples remaining denoising errors along reliable 3D tracks without suppressing view-dependent appearance.
Compared with prior baselines, VGGT-Diff synthesizes novel views that remain more faithful to the real scene, better preserving its original geometry, object identity, and visual content. The predicted images also align more accurately with the requested camera viewpoints and their corresponding ground-truth views, producing more consistent spatial layouts and fewer viewpoint or content mismatches.
Twenty full-resolution 80-frame continuous-view comparisons generated from six input views. Every sequence preserves its original GT-left / VGGT-Diff Prediction-right layout and embedded label bar.
Lightweight previews play automatically · Select a scene for the original resolution@article{chen2026vggtdiff,
title={VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis},
author={Chen, Kangjie and Li, Xiangyu and Zhang, Dongbin and Zheng, Chaoda and Chen, Shijia and Deng, Jinhao and Lin, Hongbin and Choo, Sin Wai and Wang, Minqi and Yang, Minghao and Zhong, Dake and Song, Guorui and Zhang, Yu and Liu, Xianming and Wang, Boyang},
journal={arXiv preprint arXiv:2609.33253},
year={2026}
}