Speaker

Talk

YouTube link

Ankit Goyal

Ankit Goyal
Research Scientist
NVIDIA
View Transformers for 3D Manipulation in Robotics
Thu, Sep 26, 2024 · 11:30 AM ET
Abstract

In this presentation, we introduce our work on View Transformers that achieve state-of-the-art results in 3D object manipulation: RVT and its successor, RVT-2. RVT is a multi-view transformer model designed to handle complex 3D manipulation tasks by predicting gripper pose actions using camera images and task descriptions. It demonstrated significant improvements over existing methods like PerAct, achieving 26% higher task success and training 36 times faster, while also providing real-world applicability with few demonstrations. Building on this foundation, RVT-2 addresses challenges in high-precision tasks, offering 6 times faster training and doubling inference speed compared to RVT. RVT-2 sets a new benchmark on the RLBench dataset, improving task success from 65% to 82%, and shows strong real-world performance, particularly in tasks requiring precision.

Luca Weihs

Generalizable robotic agents through large-scale simulation
Thu, Oct 3, 2024 · 11:30 AM ET
Abstract

We show that imitating shortest-path planners in simulation produces Stretch RE-1 robotic agents that, given language instructions, can proficiently navigate, explore, and manipulate objects in both simulation and in the real world using only RGB sensors (no depth maps or GPS coordinates). This surprising result is enabled by our end-to-end, transformer-based, SPOC architecture, powerful visual encoders paired with extensive image augmentation, and the dramatic scale and diversity of our training data: millions of frames of shortest-path-expert trajectories collected inside approximately 200,000 procedurally generated houses containing 40,000 unique 3D assets. We further show how scale and careful caching enable online reinforcement learning with transformers in these diverse environments and allow us to produce even more masterful agents. All data, code, and pretrained models are open-source and available online.

Mandi Zhao

New Perspectives on Harnessing Foundation Models for Robot Learning
Thu, Oct 10, 2024 · 11:30 AM ET
Abstract

Foundation models are both appealing and perplexing to roboticists: they possess clear superior capabilities in reasoning and understanding our world that should be useful for robots; however, the high-level pre-trained capabilities are often lacking at low-level embodied tasks, hence it’s often unclear how to effectively incorporate these models into robotic systems. In this talk, I will present two recent projects that explore different approaches to this problem, and hope the lessons and discussions here will spark new ideas and direction in the future. First is RoCo: Dialectic Multi-Robot Collaboration with Large Language Models (https://arxiv.org/abs/2307.04738), which uses zero-shot Large Language Models (LLMs) as a communication tool to facilitate multi-robot collaboration. Second is Real2Code: Reconstruct Articulated Objects via Code Generation (https://arxiv.org/abs/2406.08474), where we adapt both LLM and large pre-trained vision models to propose a new approach to the Real2Sim2Real problem for articulated objects.

Lukas Schmid

Lukas Schmid
Postdoctoral Fellow
MIT
Active Perception and Long-term Understanding of Dynamic Human-Centric Scenes
Thu, Oct 24, 2024 · 11:30 AM ET
Abstract

The ability to build an actionable representation of the environment of a robot is crucial for autonomy and prerequisite to a large variety of applications, ranging from home, service, care and consumer robots to industrial inspection and maintenance to disaster response and defense missions. Notably, a large part of the promise of autonomous robots depends on long-term operation in domains shared with humans and other agents. These environments are typically highly complex, semantically rich, and immensely dynamic with agents regularly moving through and interacting with the scene.

This talk presents an autonomy pipeline to address these challenges. We first present methods for detection and representation complex semantics, short-term motion, and long-term changes in dense robot maps. We then show how these challenges can be addressed jointly and in real-time in a unified framework we call Khronos. We show how Dynamic Scene Graphs (DSGs) can represent semantic symbols in a task-driven fashion and facilitate the prediction of likely future evolutions of the scene based on the data the robot has already collected. Lastly, we show how robots as embodied agents can leverage our actionable scene representations and predictions to complete tasks such as actively gathering data that helps them improve their world models and perception capabilities over time. The presented methods are demonstrated on-board fully autonomous aerial and ground robots, run in real-time on the limited hardware available, and are available as open-source software.

Vincent Leroy

Vincent Leroy
Research Scientist
NAVER LABS Europe
From CroCo to MASt3R: A Paradigm Change in 3D Vision
Thu, Nov 7, 2024 · 11:30 AM ET
Abstract

In the rapidly evolving field of 3D vision, a multitude of heterogeneous downstream tasks coexist, such as visual localization, depth estimation, 3d reconstruction, etc. The current paradigm is to develop a dedicated method to solve each task, thereby largely ignoring their potential inter-connections and synergies. Developing unified models able to handle multiple 3D geometric downstream tasks remains a challenge. This presentation introduces three interconnected advancements in this respect: CroCo, DUSt3R and MASt3R. CroCo, a self-supervised pre-training framework, utilizes a pretext task to lay foundations for DUSt3R/MASt3R, a unified foundational model for geometric 3D vision. Cross-view completion, or CroCo in short, first serves to learn robust representations of 3D geometry from pairs of images depicting the same scene from different viewpoints. By masking parts of an image and predicting these given another viewpoint of the scene, CroCo effectively captures priors about spatial relationships and geometric information, setting a strong foundation for downstream 3D vision tasks. Then, building on the robust pre-trained models provided by CroCo, DUSt3R introduces a novel approach for Dense Unconstrained Stereo 3D Reconstruction. This method revolutionizes traditional multi-view stereo reconstruction by regressing pointmaps that encode scene geometry without requiring calibrated or posed cameras. DUSt3R simplifies the complex pipeline of traditional 3D reconstruction methods, significantly reducing computational overhead and enhancing performance across various benchmarks. MASt3R further extends DUSt3R by adding the ability to establish accurate pixel correspondences. The journey from CroCo to MASt3R exemplify a significant paradigm shift in 3D vision technologies. This presentation will delve into the methodologies, innovations, and synergistic integration of these frameworks, demonstrating their impact on the field and potential future directions. The discussion aims to highlight how these advancements unify and streamline the processing of 3D visual data, offering new perspectives and capabilities in robotic navigation, cultural heritage preservation, and beyond.

Michael Yip

Michael Yip
Associate Professor
UC San Diego
The Language of Learning-based Robot Motion Planning
Thu, Nov 14, 2024 · 11:30 AM ET
Abstract

Robots and other autonomous systems need to understand how to move in complex and dynamic environments while avoiding or minimizing unwanted contact. With over 40 years of evolution, classical motion planning solutions have been hitting practical limits in solving many real-world environments due to their unpredictability as well as the curse-of-dimensionality. Even with today's best algorithms, we often experience unsatisfactory behaviors or performance: with robots taking many seconds or even minutes to think before they move, and even then, the movement may appear unusually roundabout and suboptimal. Higher-level considerations, including safety, responsiveness, and accounting for uncertainty can also add significant challenges.

Now, Machine Learning has arrived to solve the motion planning problem and promises to provide a transformative leap in autonomy. How does it manage to achieve this? In this talk, I will introduce our work in motion planning networks that started this path toward neural planners, breaking the mold of how robots should plan for navigation. In both simulation and real-world examples, we show how this research area has grown to solve multi-manipulator coordination, task and motion planning, constrained behaviors, as well the visual perception needed to make it all work in everyday environments.

Bryan Chan

Why can’t we use reinforcement learning for image-based robotic manipulation?
Thu, Nov 21, 2024 · 11:30 AM ET
Abstract

Imitation learning (IL) is a popular method for robot learning due to the wider data availability, improved data-collection techniques, and the development of vision-language models. However IL can only do as well as the demonstrators. In this case, reinforcement learning (RL) is especially compelling as RL offers an autonomous self-improvement cycle. However, current RL algorithms seemingly fail to learn efficiently in image-based robotic manipulation tasks. In this talk we discuss the potential challenges of RL in robot learning and we present how we might be able to resolve some of these challenges.

Sayantan Auddy

Continual Learning for Robotic Manipulation
Thu, Nov 28, 2024 · 11:30 AM ET
Abstract

Humans continuously acquire new knowledge while retaining and refining what they have already learned. As robotics advances, continual learning is poised to become a critical feature, enabling robots to navigate the ever-changing demands of real-world environments. In this talk, I will discuss my work on continual learning of manipulation skills in real-world scenarios. Based on the nature of the continually learned tasks, my talk is divided into two parts.

The first part addresses scenarios in which distinct manipulation tasks with different objectives need to be learned, such as opening a box or pouring from a cup. Here I will discuss "Continual Learning from Demonstration" (continual LfD), a method based on hypernetworks and neural ODEs that is used to learn a sequence of manipulation tasks from human demonstrations. This method effectively retains multiple LfD skills without storing or retraining on past demonstrations and requires only a few demonstrations per task. Next, I will discuss an approach to "stable" continual LfD, where we show that non-divergent, stable motion prediction improves continual learning performance and makes it possible to scale to a higher number of learned tasks and remember LfD tasks in high-dimensional spaces.

The second part of my talk covers scenarios where the robot learns the same basic manipulation task but under varying environmental conditions. I will discuss "Continual Domain Randomization", a method that combines continual learning with domain randomization to achieve sim-to-real transfer. Here, a robot is trained in a sequence of differently randomized simulated environments and utilizes regularization-based continual learning to remember the effect of each environment. This results in a trained agent that can be directly transferred to the physical robot and exhibits robust zero-shot performance.