Speaker

Talk

YouTube link

Iro Armeni

Iro Armeni
Assistant Professor
Stanford
Rectified Point Flow: Generic Point Cloud Pose Estimation
Thu, Jan 22, 2026 · 11:00 AM ET
Abstract

Robotic systems frequently confront geometric alignment problems such as point cloud registration and multi-part object assembly, which are typically addressed with task-specific pipelines and explicit pose regression. In this talk, I will present Rectified Point Flow, a unified formulation that casts both problems as a single conditional generative task by learning a continuous point-wise velocity field that transports unposed points to their target locations. This approach naturally recovers part poses and intrinsically captures object symmetries without supervision, outperforming prior methods across six benchmarks. I will conclude by discussing applications in cultural heritage, where robust alignment and assembly are critical for documenting, restoring, and reassembling fragmented artifacts and monuments.

Majid Khadiv

Majid Khadiv
Assistant Professor
TUM
A Scalable Path Towards Humanoid Foundation Models
Thu, Jan 29, 2026 · 11:00 AM ET
Abstract

Recent advances in foundation models have shown great promise in imitating teleoperation demonstrations for complex manipulation tasks. While highly successful, current methods rely heavily on large amounts of teleoperation data, which does not scale to building a humanoid foundation model. In this talk, I outline two key developments that together provide a scalable framework for generating the large-scale data needed to train humanoid foundation models. The first is a general optimization-based task and motion planning (TAMP) framework that generates diverse strategies for accomplishing different tasks, and can also leverage pre-trained VLMs to propose manipulation sequences in the form of subgoals. The second focuses on retargeting the vast amounts of available human motion data to humanoid robots. I will demonstrate how these data sources can be leveraged to build a foundation model for humanoid robots. I will conclude with a brief overview of our work on safety.

Yunzhu Li

Yunzhu Li
Assistant Professor
Columbia
Scaling Robotic Manipulation via Structured World Models and Tactile Sensing
Thu, Feb 5, 2026 · 11:00 AM ET
Abstract

Scaling robotic manipulation demands both predictive models of how the world evolves under action and rich sensing of physical contact. In this talk, I will present two complementary research directions that address these challenges. First, inspired by human intuitive physics, I introduce structured world models that incorporate physical priors through particle- and graph-based neural dynamics. These models enable model-based planning and control across a wide range of rigid, deformable, articulated, and granular objects, and support long-horizon, contact-rich manipulation. They also facilitate the construction of neural and physics-informed digital twins for scalable data generation, policy iteration, and evaluation. Second, I will present our work on scalable tactile sensing, from uncovering principles of human grasping with dense tactile gloves to developing flexible, low-cost tactile arrays and portable visuo-tactile grippers. Combined with simulation and large-scale real-world data collection, these tactile systems enable robust learning and improved sim-to-real transfer for tasks involving visual occlusion, fragile objects, and complex physical interactions. Together, these efforts highlight key ingredients for scaling robotic manipulation toward greater generality, robustness, and physical competence, laying the groundwork for physically grounded foundational robotic models.

Andrew Wagenmaker

Andrew Wagenmaker
Postdoctoral Scholar
UC Berkeley
What Does RL Theory Have to Do with Robotics?
Thu, Feb 12, 2026 · 11:00 AM ET
Abstract

While the theory of reinforcement learning has advanced to a fairly mature place, it is often not apparent how this theory can impact practice. This is especially true in domains such as robotics, where the challenges faced by practitioners typically feel far removed from the settings and algorithms considered by theorists. In this talk, I will discuss how RL theory can impact practice in robotics despite this apparent gap. I will focus in particular on two case studies centered around the question of pretraining for online adaptation. In the first case, I will explore the question of sim-to-real transfer for robotics, and how we should pretrain with RL in a simulator to enable effective transfer to the real world. In the second case, I will discuss how we can pretrain a policy from human demonstration data to ensure it is a good initialization for further RL finetuning. In both cases, I will show how theory provides the key algorithmic insights that lead to highly effective practical approaches that enable real-world robot learning.

Makram Chahine

Makram Chahine
PhD Student
MIT
From Internet to Edge and Back
Thu, Feb 19, 2026 · 11:00 AM ET
Abstract

Modern Robotics faces a central dilemma: unprecedented access to internet-scale models for digital perception and reasoning, contrasted with uncompromising constraints of the physical world. The challenge lies in bridging these high-level capabilities with real-time edge compute, limited connectivity, and safety-critical demands, all while operating on demonstration datasets that are orders of magnitude smaller than the trillions of tokens fueling frontier foundation models. This talk explores both ends of the efficiency spectrum: first, through decentralized algorithms that leverage foundation models to enable high-stakes swarm wildlife monitoring, and second, through in-training architectural compression techniques to make large models viable on the edge. First, we go from internet to edge: bringing state-of-the-art vision models onto a swarm of drones for autonomous sperm whale monitoring as part of Project CETI. We present a fully decentralized pipeline covering scouting, detection, motion planning, multi-agent registration, goal assignment, and individual monitoring execution.Next, we go from edge back to training: asking whether large models can be reduced during training itself. We introduce CompreSSM, a control-theoretic framework that leverages Hankel singular values and balanced truncation to progressively compress State Space Models while they learn, with gains on both training and inference computational costs.Together, the two works trace a full loop: digital intelligence deployed on physical robots, and physical resource constraints feeding back to reshape how we train models in the first place.

Nils Dengler

Learning Perception and Manipulation of Objects in Cluttered Environments
Thu, Mar 5, 2026 · 11:00 AM ET
Abstract

Autonomous robots operating in unstructured real-world environments must perceive partially observable scenes, reason under uncertainty, and manipulate objects while respecting physical constraints. This talk presents model-free and uncertainty-aware approaches that integrate non-prehensile manipulation, probabilistic semantic mapping, and uncertainty-driven action selection. First, I will introduce learning-based pushing strategies that enable robots to interact with and rearrange objects directly from experience. Building on this, I will present a probabilistic semantic mapping framework that integrates manipulation to actively resolve occlusions and improve scene understanding. Finally, I will show how reinforcement learning can leverage these uncertainty-aware representations to select perception and manipulation actions that efficiently reduce uncertainty while minimizing unnecessary disturbance to the environment.

Anirudha Majumdar

Anirudha Majumdar
Associate Professor
Princeton University
Visiting Research Scientist
Google DeepMind
Trustworthy World Models for Safe Generalist Robots
Thu, Mar 12, 2026 · 11:00 AM ET
Abstract

Action-conditioned video generation models have the potential to serve as general-purpose world models for robotics. Their ability to generate photorealistic observations, simulate complex physical interactions, and be improved with data make them an attractive alternative to traditional physics-based models for policy evaluation, reinforcement learning, and inference-time planning. In this talk, I will highlight our recent work from Google DeepMind on using video models as "simulators" for evaluating generalist robot policies for performance, generalization, and safety. I will then highlight challenges with hallucinations in current video models: objects can appear or disappear, deform in unrealistic ways, or move in a manner that defies physics. I will describe work from my group at Princeton on addressing these challenges using autonomous play data, and demonstrate how the resulting models lead to significant improvements for policy evaluation and reinforcement learning inside the world model. Finally, I will talk about world models that know when they don’t know through rigorous uncertainty quantification; to our knowledge, this is the first work on calibrated uncertainty quantification for action-conditioned video models.

Christina Kassab

Christina Kassab
Ph.D. Student
Oxford Robotics Institute
University of Oxford
A Unified Approach to Semantic and Geometric Mapping
Thu, Apr 2, 2026 · 11:00 AM ET
Abstract

Visual SLAM has long provided the geometric backbone for autonomous systems, yet traditional systems lack the semantic richness needed for robots to interact meaningfully with their surroundings. This talk explores how vision-language models can bridge this gap, addressing two central questions: how beneficial are such models for SLAM tasks, and how much can a single pre-trained model accomplish across multiple scene understanding tasks? I will present LEXIS, a real-time semantic SLAM system that uses CLIP for open-vocabulary room classification, place recognition, and loop closure within a unified framework. I will then examine the challenges of open-vocabulary 3D object segmentation, and introduce OpenLex3D, a tiered benchmark that evaluates these systems beyond closed-vocabulary metrics. Finally, I will present ongoing work on LEXI-SG which combines semantic capabilities with feed-forward reconstruction models to build open-vocabulary scene graphs from monocular video.

Tom Silver

Tom Silver
Assistant Professor
Princeton University
Towards Robots that Learn In-The-Wild by Engineering Their Own Software
Thu, Apr 9, 2026 · 11:00 AM ET
Abstract

Every user, environment, and situation is unique, and robots should adapt accordingly. This kind of in-the-wild learning is fundamentally different from the large-scale pretraining that happens before deployment. Robots in the wild must learn a lot from a little, stay transparent even as their behavior changes, and never compromise safety. In this talk, I will argue that software engineering offers a compelling paradigm for this kind of learning. Code is composable, hierarchical, interpretable, testable, debuggable, and extensible—all properties we want for robots that are changing on-the-fly. The central challenge is that robot code is necessarily a lossy abstraction of the physical world. This boundary between world and code is where the hardest and most interesting problems arise. I will describe our work at this boundary that combines neuro-symbolic learning, program synthesis, and task and motion planning to create robots that act as self-supervised software engineers in-the-wild.

Coline Devin

Coline Devin
Member of Technical Staff
Generalist AI
Scaling Foundation Models for Robot Manipulation
Thu, Apr 16, 2026 · 11:00 AM ET

Angelica Lim

Multimodal and Socially Interactive Embodied AI
Thu, Apr 23, 2026 · 11:00 AM ET
Abstract

The long-term vision of socially interactive humanoid robots requires machines that can engage with humans through their bodies, adapting in real time to a partner's movement, intent, and ability level. Like an embodied ChatGPT, this problem requires generating appropriate motion responses to an active participant whose behavior cannot be scripted or predicted; further, the responses must carry social meaning, appropriate to the context, the partner, and the shared interaction. In this talk, we will explore human intent understanding and humanoid motion generation across multiple nonverbal dynamic modalities, including face, gesture, prosody and trajectory. I'll discuss how our lab has investigated the nebulous concept of “context’ and explore how it modulates the way that humans express themselves, as well as how it changes how expressions are perceived. We focus especially on generative models and tasks with subjective evaluations, towards robots that are acceptable to humans in society.

Steven Parkison

Robot Navigation for Inspection and Intervention
Thu, Apr 30, 2026 · 11:00 AM ET
Abstract

As infrastructure assets age, the required cadence of inspection and maintenance tasks increases to certify their continued safe and efficient operation. These tasks traditionally involve manual intervention by skilled technicians in remote sites. Robotic systems offer a way of increasing the productivity of these technicians while simultaneously removing them from dangerous situations. This talk will discuss how modern AI and robotic perception methods can be leveraged to enable autonomy in these challenging environments, including inside the penstocks of hydroelectric power plants, along high-voltage transmission lines, and within remote power substations.