Stochastic Interpolants for Controllable Scene Generation

Stochastic interpolants offer a powerful toolbox for generative modeling, including denoising diffusion, flow matching, and inductive moment matching. In the context of ETHAR, they hold great promise as a path towards controllable sampling and synthesis of multimodal content across the 2D and 3D domains. This potential is exemplified by scenarios where users aim to reimagine or enhance physical environments based on minimal imagery and high-level guidance. The proposed framework leverages stochastic interpolants to generate semantically consistent and geometrically plausible 3D reconstructions, including novel views that complete and enrich the scene beyond what is directly observable. We propose to develop a framework for interactive, semantically guided, and geometrically coherent content creation based on stochastic interpolants with variational regularization, in service of immersive user experiences and augmented reality.

Project Team

Prof. Konrad Schindler
ETH Zurich
Dr. Dominik Narnhofer
ETH Zurich
PD Dr. Federico Tombari
Google

Publications

Understanding, Accelerating, and Improving MeanFlow Training

Authors:
Jin-Young Kim
Hyojun Go
Lea Bogensperger
Julius Erbach
Nikolai Kalischek
PD Dr. Federico Tombari
Prof. Konrad Schindler
Dr. Dominik Narnhofer
Published:
to be published at CVPR, 2026

Generative, panoramic depth estimation (working title)

Published:
submitted to NeurIPS, 2026

VIST3A: Text-to-3D by Stitching a Multi-View Reconstruction Network to a Video Generator

Authors:
Hyojun Go
Dr. Dominik Narnhofer
Goutam Bhat
Prune Truong
PD Dr. Federico Tombari
Prof. Konrad Schindler
Published:
ICLR 2026 (Oral)