Learning Human Interaction from Videos

This project aims to develop self-supervised and unsupervised methods to learn actions directly from large-scale videos. Unlike current approaches that rely on heavy supervision or predefined latent structures, our goal is to discover an implicit latent action space with broader generalization. Such a representation can transfer across domains and support robotics, human motion and interactions, customizable video generation, and planning tasks via world models.

Project Team

Prof. Siyu Tang

Siyu Tang leads the Computer Vision and Learning Group (VLG) at the Institute of Visual Computing. Her research focuses on computer vision and machine learning, specializing in perceiving and modeling humans. Her group studies computational models that enable machines to perceive and analyze human pose, motion, and activities from visual input. The group leverages machine learning and optimization techniques to build statistical models of humans and their behaviors. Their goal is to advance algorithmic foundations of scalable and reliable human digitalization, enabling a broad class of real-world applications.

Prof. Bernt Schiele

Dr. Thabo Beeler

Thabo Beeler is currently a Research Director at Google, where he is heading the Syntec team within AR Perception. They work on digital humans in the context of virtual and augmented reality, focusing on capture, reconstruction, appearance acquisition, generative modeling, and synthesis. Prior to that, Thabo Beeler was a research scientist at Disney Research | Studios where he built up the Capture and Effects group, focusing on digital humans for film.

Dr. Vassilis Choutas

Dr. Korrawe Karunratanakul

Dr. Jan Eric Lenssen

Bahri Batuhan Bilecen

Publications

InvAct: View and Scene-invariant Atomic Action Learning from Videos

Published:
submitted to NeurIPS, 2026

The project is about learning atomic-level actions from human demonstration videos, in different camera viewpoint and diverse scene settings. We then demonstrate that this learnt representation can be effectively used for short and long-sequence action retrieval, action classification, and VLM-based robotic policy pretraining.