Scalable Contextualized Scene Representation for Robotics

We are building a framework for large-scale, semantically-aware Gaussian splats to be used in robotic planning. By leveraging scene graphs and shared Gaussian embeddings, our system allows for dynamic, task-driven Level-of-Detail filtering and efficient querying by Vision-Language Models. This approach will be deployed on an ANYmal legged robot in challenging outdoor environments, with a XR pilot interacting with the semantic scene.

Project Team

Prof. Marco Hutter

Marco Hutter is a Professor of Robotic Systems and the Head of the Center for Robotics at ETH Zurich. His research interests are focused on the development of novel machines and their intelligence for use in harsh and demanding environments. Together with his team, he has developed a range of walking robots, mobile manipulators, and autonomous excavators utilized in industrial inspection, construction and forestry, as household aids, and even for extraterrestrial research. Additionally, Marco is a co-founder of several ETH startups, such as ANYbotics AG and Gravis Robotics AG, which specialize in marketing legged robots and autonomous construction equipment.

Dr. Vaishakh Patil

Dr. Keisuke Tateno

PD Dr. Federico Tombari

Federico Tombari is a Research Director at Google Zurich, Switzerland, where he leads an applied research team in Computer Vision and Machine Learning across the US, Switzerland, and Germany. He is also affiliated with the Faculty of Computer Science at TUM as a lecturer (PrivatDozent) at the CAMP Chair. An up-to-date list of publications is available at his Google Scholar. Federico Tombari's research activity is focused on different aspects of computer vision and machine learning, with a focus on 3D computer vision (e.g. scene understanding, 3D object recognition, 3D reconstruction, SLAM). The fields of application of his research are mainly in robotics, augmented reality, autonomous driving and healthcare. He is currently particularly excited about unsupervised learning for visual data, Large Multimodal Models, (Neural) Radiance Fields, and scene graphs for scene understanding.

Maximum Wilder-Smith

Publications

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

Authors:
Arjun Bhardwaj
Maximum Wilder-Smith
Mayank Mittal
Dr. Vaishakh Patil
Prof. Marco Hutter
Published:
Proceedings of Robotics: Science and Systems, 2026

MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments

Authors:
Svitlana Morkva
Maximum Wilder-Smith
Michael Oechsle
Alessio Tonioni
Prof. Marco Hutter
Dr. Vaishakh Patil
Published:
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2026

DiskChunGS: Large-Scale 3D Gaussian SLAM Through Chunk-Based Memory Management

Authors:
Casimir Feldmann
Maximum Wilder-Smith
Dr. Vaishakh Patil
Michael Oechsle
Michael Niemeyer
Dr. Keisuke Tateno
Prof. Marco Hutter
Published:
IEEE Robotics and Automation Letters, 2026