Zero-Shot Novel Depth Synthesis Using 3D Foundation Model Scene Representations

Abstract

We present Z3D, a zero-shot framework for novel-view depth synthesis that leverages scene representations from 3D foundation models. Given source views and the relative pose of a target view, Z3D synthesizes dense depth for the target view without requiring target-view depth supervision.

Publication
European Conference on Computer Vision 2026

Zero-Shot Novel Depth Synthesis Using 3D Foundation Model Scene Representations

Z3D is a zero-shot framework for novel-view depth synthesis that combines 3D foundation model scene representations with a diffusion model to predict depth from sparse source views.

Given source images and the relative pose of a target view, Z3D first extracts a scene representation from the source views using a pretrained 3D foundation model. The resulting representation is then used to synthesize the depth of the target view without requiring target-view depth supervision.

Project Page · Code

Denis Mbey Akola
Denis Mbey Akola
Graduate student

My research interests include computer vision and robotics