
We present Z3D, a zero-shot framework for novel-view depth synthesis that leverages scene representations from 3D foundation models. Given source views and the relative pose of a target view, Z3D synthesizes dense depth for the target view without requiring target-view depth supervision.
Z3D is a zero-shot framework for novel-view depth synthesis that combines 3D foundation model scene representations with a diffusion model to predict depth from sparse source views.
Given source images and the relative pose of a target view, Z3D first extracts a scene representation from the source views using a pretrained 3D foundation model. The resulting representation is then used to synthesize the depth of the target view without requiring target-view depth supervision.