Scalable Contextualized Scene Representation for Robotics

We are building a framework for large-scale, semantically-aware Gaussian splats to be used in robotic planning. By leveraging scene graphs and shared Gaussian embeddings, our system allows for dynamic, task-driven Level-of-Detail filtering and efficient querying by Vision-Language Models. This approach will be deployed on an ANYmal legged robot in challenging outdoor environments, with a XR pilot interacting with the semantic scene.

Project Team

Prof. Marco Hutter
ETH Zurich
Dr. Vaishakh Patil
ETH Zurich
Dr. Keisuke Tateno
Google
PD Dr. Federico Tombari
Google
Maximum Wilder-Smith
ETH Zurich

Publications

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

Authors:
Arjun Bhardwaj
Maximum Wilder-Smith
Mayank Mittal
Dr. Vaishakh Patil
Prof. Marco Hutter
Published:
Proceedings of Robotics: Science and Systems, 2026

MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments

Authors:
Svitlana Morkva
Maximum Wilder-Smith
Michael Oechsle
Alessio Tonioni
Prof. Marco Hutter
Dr. Vaishakh Patil
Published:
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2026

DiskChunGS: Large-Scale 3D Gaussian SLAM Through Chunk-Based Memory Management

Authors:
Casimir Feldmann
Maximum Wilder-Smith
Dr. Vaishakh Patil
Michael Oechsle
Michael Niemeyer
Dr. Keisuke Tateno
Prof. Marco Hutter
Published:
IEEE Robotics and Automation Letters, 2026