Fall 2024
|
|
Speaker |
Talk |
YouTube link |
|
|
View Transformers for 3D Manipulation in Robotics
Thu, Sep 26, 2024 · 11:30 AM ET AbstractIn this presentation, we introduce our work on View Transformers that achieve state-of-the-art results in 3D object manipulation: RVT and its successor, RVT-2. RVT is a multi-view transformer model designed to handle complex 3D manipulation tasks by predicting gripper pose actions using camera images and task descriptions. It demonstrated significant improvements over existing methods like PerAct, achieving 26% higher task success and training 36 times faster, while also providing real-world applicability with few demonstrations. Building on this foundation, RVT-2 addresses challenges in high-precision tasks, offering 6 times faster training and doubling inference speed compared to RVT. RVT-2 sets a new benchmark on the RLBench dataset, improving task success from 65% to 82%, and shows strong real-world performance, particularly in tasks requiring precision. |
||
|
|
Generalizable robotic agents through large-scale simulation
Thu, Oct 3, 2024 · 11:30 AM ET AbstractWe show that imitating shortest-path planners in simulation produces Stretch RE-1 robotic agents that, given language instructions, can proficiently navigate, explore, and manipulate objects in both simulation and in the real world using only RGB sensors (no depth maps or GPS coordinates). This surprising result is enabled by our end-to-end, transformer-based, SPOC architecture, powerful visual encoders paired with extensive image augmentation, and the dramatic scale and diversity of our training data: millions of frames of shortest-path-expert trajectories collected inside approximately 200,000 procedurally generated houses containing 40,000 unique 3D assets. We further show how scale and careful caching enable online reinforcement learning with transformers in these diverse environments and allow us to produce even more masterful agents. All data, code, and pretrained models are open-source and available online. |
||
|
|
New Perspectives on Harnessing Foundation Models for Robot Learning
Thu, Oct 10, 2024 · 11:30 AM ET AbstractFoundation models are both appealing and perplexing to roboticists: they possess clear superior capabilities in reasoning and understanding our world that should be useful for robots; however, the high-level pre-trained capabilities are often lacking at low-level embodied tasks, hence it’s often unclear how to effectively incorporate these models into robotic systems. In this talk, I will present two recent projects that explore different approaches to this problem, and hope the lessons and discussions here will spark new ideas and direction in the future. First is RoCo: Dialectic Multi-Robot Collaboration with Large Language Models (https://arxiv.org/abs/2307.04738), which uses zero-shot Large Language Models (LLMs) as a communication tool to facilitate multi-robot collaboration. Second is Real2Code: Reconstruct Articulated Objects via Code Generation (https://arxiv.org/abs/2406.08474), where we adapt both LLM and large pre-trained vision models to propose a new approach to the Real2Sim2Real problem for articulated objects. |
||
|
|
Active Perception and Long-term Understanding of Dynamic Human-Centric Scenes
Thu, Oct 24, 2024 · 11:30 AM ET AbstractThe ability to build an actionable representation of the environment of a robot is crucial for autonomy and prerequisite to a large variety of applications, ranging from home, service, care and consumer robots to industrial inspection and maintenance to disaster response and defense missions. Notably, a large part of the promise of autonomous robots depends on long-term operation in domains shared with humans and other agents. These environments are typically highly complex, semantically rich, and immensely dynamic with agents regularly moving through and interacting with the scene. |
||
|
|
From CroCo to MASt3R: A Paradigm Change in 3D Vision
Thu, Nov 7, 2024 · 11:30 AM ET AbstractIn the rapidly evolving field of 3D vision, a multitude of heterogeneous downstream tasks coexist, such as visual localization, depth estimation, 3d reconstruction, etc. The current paradigm is to develop a dedicated method to solve each task, thereby largely ignoring their potential inter-connections and synergies. Developing unified models able to handle multiple 3D geometric downstream tasks remains a challenge. This presentation introduces three interconnected advancements in this respect: CroCo, DUSt3R and MASt3R. CroCo, a self-supervised pre-training framework, utilizes a pretext task to lay foundations for DUSt3R/MASt3R, a unified foundational model for geometric 3D vision. Cross-view completion, or CroCo in short, first serves to learn robust representations of 3D geometry from pairs of images depicting the same scene from different viewpoints. By masking parts of an image and predicting these given another viewpoint of the scene, CroCo effectively captures priors about spatial relationships and geometric information, setting a strong foundation for downstream 3D vision tasks. Then, building on the robust pre-trained models provided by CroCo, DUSt3R introduces a novel approach for Dense Unconstrained Stereo 3D Reconstruction. This method revolutionizes traditional multi-view stereo reconstruction by regressing pointmaps that encode scene geometry without requiring calibrated or posed cameras. DUSt3R simplifies the complex pipeline of traditional 3D reconstruction methods, significantly reducing computational overhead and enhancing performance across various benchmarks. MASt3R further extends DUSt3R by adding the ability to establish accurate pixel correspondences. The journey from CroCo to MASt3R exemplify a significant paradigm shift in 3D vision technologies. This presentation will delve into the methodologies, innovations, and synergistic integration of these frameworks, demonstrating their impact on the field and potential future directions. The discussion aims to highlight how these advancements unify and streamline the processing of 3D visual data, offering new perspectives and capabilities in robotic navigation, cultural heritage preservation, and beyond. |
||
|
|
The Language of Learning-based Robot Motion Planning
Thu, Nov 14, 2024 · 11:30 AM ET AbstractRobots and other autonomous systems need to understand how to move in complex and dynamic environments while avoiding or minimizing unwanted contact. With over 40 years of evolution, classical motion planning solutions have been hitting practical limits in solving many real-world environments due to their unpredictability as well as the curse-of-dimensionality. Even with today's best algorithms, we often experience unsatisfactory behaviors or performance: with robots taking many seconds or even minutes to think before they move, and even then, the movement may appear unusually roundabout and suboptimal. Higher-level considerations, including safety, responsiveness, and accounting for uncertainty can also add significant challenges. |
||
|
|
Why can’t we use reinforcement learning for image-based robotic manipulation?
Thu, Nov 21, 2024 · 11:30 AM ET AbstractImitation learning (IL) is a popular method for robot learning due to the wider data availability, improved data-collection techniques, and the development of vision-language models. However IL can only do as well as the demonstrators. In this case, reinforcement learning (RL) is especially compelling as RL offers an autonomous self-improvement cycle. However, current RL algorithms seemingly fail to learn efficiently in image-based robotic manipulation tasks. In this talk we discuss the potential challenges of RL in robot learning and we present how we might be able to resolve some of these challenges. |
||
|
|
Continual Learning for Robotic Manipulation
Thu, Nov 28, 2024 · 11:30 AM ET AbstractHumans continuously acquire new knowledge while retaining and refining what they have already learned. As robotics advances, continual learning is poised to become a critical feature, enabling robots to navigate the ever-changing demands of real-world environments. In this talk, I will discuss my work on continual learning of manipulation skills in real-world scenarios. Based on the nature of the continually learned tasks, my talk is divided into two parts. |