Visual Geometry Meets Diffusion
for Sparse-View Novel View Synthesis

1 XPeng Motors2 The Chinese University of Hong Kong3 Tsinghua University

∗ Equal contribution† Corresponding author

Effective 21K checkpoint · Full resolution · 80 continuous frames

Full-Resolution Long-Sequence Demo

From six input images, VGGT-Diff generates a continuous 80-frame novel-view sequence at full resolution. The model was initialized from a 20K half-resolution checkpoint, then received only 1K steps of full-resolution continuation training.

1080p autoplay · 24 fps · 40 seconds · sound on when browser permits
Watch in 4KDownload 4K
Training status.

This effective 21K checkpoint is still training. Fine details and cross-view consistency are expected to continue improving. Fully trained full-resolution checkpoints will be released on the VGGT-Diff GitHub project page.

Six Full-Resolution Input Views

832 × 480 each
Full-resolution input view 01
Input 01
Full-resolution input view 02
Input 02
Full-resolution input view 03
Input 03
Full-resolution input view 04
Input 04
Full-resolution input view 05
Input 05
Full-resolution input view 06
Input 06

Abstract

We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing NVS methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model.

A confidence-aware Visual Geometry Router transforms 3D points and appearance-bearing tokens into query-aligned conditions while preserving front and back surface evidence. Joint target-view denoising is regularized by Point-Track Residual Consistency, which aligns predicted-clean residuals along reliable 3D tracks. Uncertainty-aware training and inference further improve robustness to imperfect geometry.

Method

VGGT-Diff method pipeline
Sparse input views are encoded into clean appearance latents, camera rays, and routed visual geometry. A Wan-initialized multi-view DiT jointly denoises one to sixteen requested views; flow matching and PTRC supervise fidelity and cross-view consistency.
01

Visual Geometry Router

Projects VGGT-Ω features into requested cameras with confidence-aware, two-layer visibility routing.

02

Joint Multi-View Diffusion

Processes source and query slots together so requested views exchange information during denoising.

03

Point-Track Consistency

Couples remaining denoising errors along reliable 3D tracks without suppressing view-dependent appearance.

+ Visual Geometry Conditioning
Confidence-aware visual geometry routing diagram

The router begins from a sharp depth-selected anchor, preserves secondary-surface evidence near uncertain occlusions, and learns only a residual correction from layered features and routing statistics.

Results

18.104DL3DV PSNR↑ best
0.222DL3DV LPIPS↓ best
16.296Mip-NeRF 360 PSNRzero-shot
−26.7%Chamfer-L1vs. FrameCrafter

Compared with prior baselines, VGGT-Diff synthesizes novel views that remain more faithful to the real scene, better preserving its original geometry, object identity, and visual content. The predicted images also align more accurately with the requested camera viewpoints and their corresponding ground-truth views, producing more consistent spatial layouts and fewer viewpoint or content mismatches.

Ground truth for Scene 1Ground Truth
FrameCrafter result
VGGT-Diff result
VGGT-DiffFrameCrafter
SEVA result
VGGT-Diff result
VGGT-DiffSEVA
DepthSplat result
VGGT-Diff result
VGGT-DiffDepthSplat
Loading scene

Drag any divider to compare. When idle, the sliders move gently to reveal both results.

Five paper-selected DL3DV scenes. Each slider compares the same target view and common 480 × 480 evaluation crop.
+ Baseline Comparisons
DL3DV
MethodTraining ScenesPSNR ↑SSIM ↑LPIPS ↓DreamSim ↓
DepthSplat77.5K17.2840.5660.3070.158
SEVA80K16.1500.4700.2530.088
FrameCrafter1K17.1800.4450.2230.066
VGGT-Diff1K18.1040.5080.2220.062
Mip-NeRF 360 zero-shot
MethodTraining ScenesPSNR ↑SSIM ↑LPIPS ↓DreamSim ↓
DepthSplat77.5K15.9770.3700.4100.215
SEVA80K14.5900.2940.3720.137
FrameCrafter1K15.6400.2790.3650.111
VGGT-Diff1K16.2960.3170.3150.091

All methods receive the same six input views and target cameras within each benchmark. VGGT-Diff uses only 1K training scenes—the same scale as FrameCrafter and far fewer than DepthSplat (77.5K) or SEVA (80K)—then transfers to Mip-NeRF 360 without dataset-specific fine-tuning.

+ PSNR Across Pose Difficulties

We stratify 192 DL3DV targets by interpolation or extrapolation and by near, mid, or far distance from the closest input camera. This controlled comparison measures how reliably each method follows increasingly difficult target cameras rather than hiding challenging cases inside a single average.

PSNR on DL3DV across target-camera difficulty. Each bin contains 32 targets and uses independent six-to-one inference.
MethodInterpolationExtrapolation
NearMidFarNearMidFar
Regression-based models
AnySplat14.19612.41510.86213.97712.56611.827
E-RayZer17.38814.87813.82416.83314.47714.712
LVSM20.27616.59913.96920.79016.63315.579
DepthSplat20.37617.42414.82020.65317.94316.374
Diffusion-based models
Aether14.03012.16711.06515.73912.29511.795
GEN3C13.62012.35611.54814.28813.05012.288
MVSplat36015.28214.42013.38915.54214.47914.176
FrameCrafter20.33817.48514.64921.16817.88816.594
SEVA20.94017.86114.70322.31118.05016.575
VGGT-Diff20.92018.39115.85322.39618.79317.793

VGGT-Diff achieves the highest PSNR in five of six bins and is within 0.020 dB of SEVA on interpolation-near, showing robust camera control from nearby interpolation to distant extrapolation.

+ Cross-View Geometry

To test whether multiple generated views agree on a single 3D structure, we jointly synthesize eight target views from the same six input views and reconstruct them with two frozen geometry models: VGGT-Ω, our primary probe, and Pi3, an independent reconstructor never used to train or condition VGGT-Diff. Each point cloud is compared with the reference reconstructed from the corresponding ground-truth views under identical cameras and settings. Agreement under both probes shows that the gains reflect stronger cross-view geometry rather than evaluator bias.

Point-cloud consistency comparison across SEVA, FrameCrafter, VGGT-Diff, and ground-truth-view reconstruction
VGGT-Ω reconstruction
MethodChamfer-L1 ↓F-score@1% ↑F-score@2% ↑
LVSM0.23870.11160.2218
SEVA0.19440.20500.3244
FrameCrafter0.10000.30860.4793
VGGT-Diff0.07330.42310.6041
Pi3 reconstruction independent probe
MethodChamfer-L1 ↓F-score@1% ↑F-score@2% ↑
LVSM0.36100.06730.1525
SEVA0.25140.16400.2751
FrameCrafter0.14040.21400.3880
VGGT-Diff0.11740.30890.5021

Evaluated across 52 DL3DV scenes. Metrics are trajectory-normalized and should be compared within each reconstructor. Lower Chamfer-L1 and higher F-scores indicate stronger cross-view geometric consistency.

Video & Trajectory Demos

Twenty full-resolution 80-frame continuous-view comparisons generated from six input views. Every sequence preserves its original GT-left / VGGT-Diff Prediction-right layout and embedded label bar.

Lightweight previews play automatically · Select a scene for the original resolution

Citation

@article{chen2026vggtdiff,
  title={VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis},
  author={Chen, Kangjie and Li, Xiangyu and Zhang, Dongbin and Zheng, Chaoda and Chen, Shijia and Deng, Jinhao and Lin, Hongbin and Choo, Sin Wai and Wang, Minqi and Yang, Minghao and Zhong, Dake and Song, Guorui and Zhang, Yu and Liu, Xianming and Wang, Boyang},
  journal={arXiv preprint arXiv:2609.33253},
  year={2026}
}