Speaker

Talk

YouTube link

Francis Engelmann

Francis Engelmann
Postdoctoral Researcher
ETH Zurich
Visiting Researcher
Google
Towards High-Fidelity Open-Vocabulary 3D Scene Understanding
Thu, Jan 11, 2024 · 11:30 AM ET
Abstract

In this talk, I will present some of our recent work on Open-Vocabulary 3D Scene Understanding. We start with Transformer-based networks and demonstrate their use as general-purpose models for a variety of 3D scene understanding tasks, including object segmentation, human body part segmentation, and vectorized floorplan reconstruction. However, despite their impressive ability to solve diverse tasks, these fully-supervised models typically fail when applied to “in-the-wild” scenes. This motivates the necessity for open-vocabulary approaches that can operate in real-world scenarios. I will then present recent works for open-vocabulary 3D scene segmentation, making use of foundation models such as CLIP and SAM. Even though the use of foundation models is revolutionizing the field in an exciting way, we will also look at current limitations and open challenges.

Yunzhu Li

Learning Structured World Models From and For Physical Interactions
Thu, Jan 25, 2024 · 11:30 AM ET
Abstract

Humans have a strong intuitive understanding of the physical world. Through observations and interactions with the environment, we build a mental model that predicts how the world would change if we applied a specific action (i.e., intuitive physics). My research draws on insights from humans and develops model-based reinforcement learning (RL) agents that learn from their interactions and build predictive models that generalize widely across a range of objects made with different materials. The core idea behind my research is to introduce novel representations and integrate structural priors into the learning systems to model the dynamics at different levels of abstraction. I will discuss how such structures can make model-based planning algorithms more effective and help robots accomplish complicated manipulation tasks (e.g., manipulating an object pile, shaping deformable foam into a target configuration, and making a dumpling from the dough using various tools). Furthermore, I will demonstrate our recent progress in learning novel scene representations, D3Fields, that are simultaneous 3D, semantic, and dynamics. I will also show how a structured scene representation can facilitate the integration of large language models (LLMs), enabling robots to perform a wide variety of everyday tasks as specified in free-form natural language.

Robotics: Enabler and Inhibitor of the SDGs
Thu, Feb 8, 2024 · 11:30 AM ET
Abstract

Robotics has the power to help our society in managing many current and foreseeable challenges, and contribute to a responsible future, as formally structured in the United Nations' Sustainable Development Goals (SDGs) initiative. In this talk, I will present our work on evaluating the impacts of robotics on the SDGs. We introduced a multidisciplinary analysis of both the enabling and disabling roles of robotics, in achieving the economic, social and environmental SDGs. We individually examined each SDG and its Targets within the context of state-of-the-art robotics documented in scientific literature. Our results indicate that robotics has the potential to enable 46 % of the Targets, particularly for the industry and environment-related SDGs, through its impacts on production, infrastructures, and monitoring activities. Inversely, robotics could inhibit 19 % of the SDG Targets, mainly through exacerbation of inequalities and tensions in the SDGs.

Xuxin Cheng

Visual Whole-Body Manipulation and Locomotion for Legged Robots
Thu, Feb 15, 2024 · 11:30 AM ET
Abstract

Legged robots are capable of traversing challenging terrains thanks to their high degrees of freedom. They are also capable of doing manipulation tasks such as pressing a button with its feet or picking and placing objects with an attached arm. However, these skills necessitate precise coordination between visual perception and whole-body movements. In this talk, I will focus on how to achieve dynamic, scalable skills for legged robots in real-world settings, by tightly combining visual feedback with whole-body movements.

Andi Peng

Andi Peng
PhD Student
MIT
Learning Abstractions from Humans
Thu, Feb 22, 2024 · 11:30 AM ET
Abstract

Generalizable robot learning is well-facilitated by abstraction—the process of creating representations that capture salient task features important for decision-making. For example, to throw away the trash—a robot may consider goal placement of trash (e.g., in the trash can) as well as movement efficiency (e.g., as fast as possible). However, these abstractions may also depend on the end user’s implicit preferences for what matters in a task (e.g. but avoid stepping on the user’s laptop), which may be difficult or infeasible for the designer to comprehensively specify beforehand. How might we capture this information safely and efficiently? This talk will explore three ways of integrating human knowledge into the abstraction-learning process: using human feedback as a general prior for creating state abstractions for imitation learning, as a personalizable interface for identifying implicit preferences, and as a pragmatic framework for learning user-aligned reward functions.

Ben Burchfiel

Ben Burchfiel
Manager, Large Behavior Models
Toyota Research Institute
Towards Large Behavior Models: Versatile and Dexterous Robots via Supervised Learning
Thu, Mar 7, 2024 · 11:30 AM ET
Abstract

Recent advances in machine learning have transformed multiple AI-related fields. Notably, robust general-purpose language and vision models are fast becoming reality and these new capabilities have already begun making their way into consumer-facing technologies where they affect the lives of many millions of people. These same underlying advancements also portend sea-change in robotics. It is now possible to reliably imbue robots with new behaviors, such as beating eggs or folding clothing, using just an hour or two of teaching and a few dozen GPU-hours of compute. In this talk, I will discuss our team's push, at the Toyota Research Institute, to scale ML-powered robot behavior teaching and the road ahead to general-purpose Large Behavior Models for robots. These models will possess the flexibility and generality of existing Large Language Models, but will be capable of dexterously controlling a robot to effect change in the physical world.

Mengdi Xu

Building Adaptable Generalist Robots
Thu, Mar 14, 2024 · 11:30 AM ET
Abstract

Deep robot learning has unlocked exciting robot capabilities in the recent decade but still struggles to generalize to unseen tasks despite large-scale pre-training. This limitation arises from the highly unstructured real world, which encompasses an endless array of possible tasks, some extending beyond the robot training set. In this talk, I will discuss building robots that can adapt at deployment with strong data efficiency, parameter efficiency, and robustness. I will highlight my research on improving robot generalization through learning to adapt under different supervisions, including in-context robot learning conditioning on a few demonstrations, unsupervised continual reinforcement learning discovering robot task structures, and embodied agents leveraging large foundation models. These approaches have shown significant promise in acquiring new motor skills in various applications and even solving long-horizon physical puzzles with creative robot tool use.

Luis Pineda
Senior Research Engineer
Meta FAIR
Robotic Dexterous Manipulation at FAIR
Thu, Mar 21, 2024 · 11:30 AM ET
Abstract

Recent advances in robotics learning and teleoperation have enabled impressive results in manipulation tasks. However, most of these have relied on overly engineered and narrow systems, and building robots with general skills such as handling novel objects or performing multiple tasks is still an open research problem. We argue that achieving general useful manipulation requires bringing together advances in perception, planning, and hardware, merging techniques from both classical robotics and purely learned methods. In this talk, we will cover recent work by our team following this paradigm. In particular, we will discuss three projects. First, I’ll present NCF-v2, a model for learning extrinsic contact that is able to improve downstream policy performance on insertion tasks. Second, I’ll present NeuralFeels, a visuo-tactile SLAM system that is able to estimate object shape and pose during in-hand manipulation in real time. Third, I’ll present RotateIt, a reinforcement learning-based method for in-hand object rotation using vision and touch. These works give a sample of the diverse work currently done at FAIR towards the problem of general dexterous manipulation, incorporating techniques from both classical and learned approaches, and showing how perception and planning advances need to come together to improve system performance and capabilities.

Lerrel Pinto

Lerrel Pinto
Assistant Professor
New York University
On Building General-Purpose Home Robots
Tue, Apr 9, 2024 · 1:00 PM ET
Abstract

The concept of a "generalist machine" in homes — a domestic assistant that can adapt and learn from our needs, all while remaining cost-effective — has long been a goal in robotics that has been steadily pursued for decades. In this talk, I will present our recent efforts towards building such capable home robots. First, I will discuss how large, pretrained vision-language models can induce strong priors for mobile manipulation tasks like pick-and-drop. But pretrained models can only take us so far. To scale beyond basic picking, we will need systems and algorithms to rapidly learn new skills. This requires creating new tools to collect data, improving representations of the visual world, and enabling trial-and-error learning during deployment. While much of the work presented focuses on two-fingered hands, I will briefly introduce learning approaches for multi-fingered hands which support more dexterous behaviors and rich touch sensing combined with vision. Finally, I will outline unsolved problems that were not obvious initially, which, when solved, will bring us closer to general-purpose home robots.

Laura Graesser
Senior Research Scientist
Google DeepMind
Practical Reinforcement Learning with Robots
Fri, Apr 12, 2024 · 1:00 PM ET
Abstract

In this talk I will discuss what the practical part of reinforcement learning means to me, deep dive into environment design in simulation and the real world, and share lessons learned about getting reinforcement learning to work on robots.

Samuele Tosatto

Efficient Action Representation for Robot Learning
Thu, Apr 18, 2024 · 11:30 AM ET

Laura C. Petrich

Becoming Bionic
Thu, Apr 25, 2024 · 11:30 AM ET
Abstract

The number of persons with limb loss is rapidly increasing, yet the ability for individuals to control smart prosthetic limbs intuitively and reliably is largely lacking. In this talk we provide a high level overview of open research problems in prosthetics as well as what we and others are currently working on to try and solve them.