Speaker

Talk

YouTube link

Jason Ma

Foundation Reward Models for General Robot Skill Acquisition
Thu, Oct 12, 2023
Abstract

General-purpose robots need to learn and adapt skills in unstructured environments. However, given the diverse nature of robotics tasks and lack of internet-scale data, the recipe of training foundation models, such as GPT-4, is difficult to replicate in robotics. Instead, I posit that foundation models for embodied intelligence should (1) provide actionable information for robots to acquire and refine skills in new environments, and (2) be capable of learning from offline non-robot data. Toward these goals, I will present my recent progress in learning foundation reward models for general robot skill acquisition from internet-scale non-robot data. First, I will show how to directly use in-the-wild human videos to supervise general-purpose value functions that can zero-shot specify dense rewards for robot tasks specified in images. Then, I will demonstrate how this value pre-training paradigm can be made general, encompassing goal specification in any modality, and in particular language. Finally, I will illustrate how these pre-trained value models can serve as an excellent subgoal decomposer, scaffolding long-horizon manipulation tasks without any human task labels.

Yilun Du and Anurag Ajay
PhD Student
MIT
PhD Student
MIT
Generative Artificial Intelligence for Decision Making
Thu, Oct 19, 2023
Abstract

Generative AI has led to stunning successes in recent years but is fundamentally limited by the amount of data available. This is especially challenging in robotics, where data is often missing and especially difficult to acquire. In this talk, we will discuss how to leverage advances in generative AI for robotics. We first introduce how video foundation models pretrained on internet data can act as a universal policy representing environments with different state and action spaces and is able to combinatorially generalize to unseen goals by using text as underlying goal specification. We then show how to compose multiple “foundation models” pre-trained on different modalities of internet data to hierarchically construct a physically executable plan to solve long-horizon robotic tasks. Finally, we illustrate the efficacy and adaptability of our approach on different long-horizon table-top manipulation tasks.

Andy Zeng

Andy Zeng
Staff Research Scientist
Google DeepMind
From words to actions
Thu, Oct 26, 2023
Abstract

The rise of recent Foundation models (and applications e.g. ChatGPT) offer an exciting glimpse into the capabilities of large deep networks trained on Internet-scale data. They hint at a possible blueprint for building generalist robot brains that can do anything, anywhere, for anyone. Nevertheless, robot data is expensive – and until we can bring robots out into the world (already) doing useful things in unstructured places, it will be challenging to match the same amount of diverse data being used to train e.g. large language models today. In this talk, I will briefly discuss some of the lessons we've learned while scaling real robot data collection, how we've been thinking about Foundation models, and how we might bootstrap off of them (and modularity) to make our robots useful sooner.

Noel Csomay-Shanklin

A Hierarchical Perspective on Robotic Control
Thu, Nov 2, 2023
Abstract

As robotic systems become increasingly capable and autonomous, both the tasks we ask them to accomplish and the resulting control strategies we develop grow in complexity. In this talk, we will discuss some perspectives on hierarchical controllers as a way to address this complexity by enabling efficient, systematic, and generalizable design processes. We will showcase recent results in achieving dynamically stable behaviors on a 3D hopping robot and a planar biped through the use of hierarchical control architectures. Along the way, we will review a quick primer on geometric and nonlinear control with a focus on how to incorporate them into the discussed hierarchical pipeline.

Vaisakh Shaj

Multi-Time Scale World Models
Thu, Nov 16, 2023
Abstract

The talk will introduce our paper "Multi Time Scale World Models" which was accepted recently in Neurips 2023 as a Spotlight. Intelligent agents use internal world models to reason and make predictions about different courses of their actions at many scales. Devising learning paradigms and architectures that allow machines to learn world models that operate at multiple levels of temporal abstractions while dealing with complex uncertainty predictions is a major technical hurdle in machine learning. In this talk, we will introduce a probabilistic formalism to learn multi-time scale world models which we call the Multi Time Scale State Space (MTS3) model. We will also discuss its computational aspects and experimental results focusing on action-conditional future predictions (dreams) spanning several seconds into the future.

Bio of Speaker: Vaisakh Shaj is pursuing his PhD in Robot Learning under Prof Gerhard Neumann at the ALR Lab, Karlsruhe Institute Of Technology, Germany. He finished his master's from the Indian Institute of Space Science and Technology. His PhD focuses on building probabilistic recurrent state space models and learning world models with it. His research interests are in developing algorithms that can learn continually in a non-stationary environment and can solve long-horizon tasks.

Dhruv Shah

Dhruv Shah
PhD Student
UC Berkeley
A General-Purpose Robotic Navigation Model
Fri, Nov 17, 2023
Abstract

The advent of large-scale machine learning models ("foundation models") has been paradigmatic in the fields of computer vision and natural language processing. What would a similar consolidation look like in robot learning, where learned models typically train on data from a single research group and on a specific robot embodiment? In this talk, I will share some recent progress we have made such a consolidation for the task of visual navigation in challenging real-world environments. I will discuss how sharing data across robots and tasks can enable remarkable generalization and adapt to new skills by a mechanism similar to soft prompting. Lastly, I will also discuss how such pre-trained models can enable exciting applications such as kilometer-scale navigation, open-vocabulary instruction following, and autonomous online improvement with reinforcement learning.

Yevgen Chebotar
Research Scientist
Google DeepMind
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Fri, Nov 17, 2023
Abstract

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot training data (such as placing an object onto a particular number or icon), and the ability to perform rudimentary reasoning in response to user commands (such as picking up the smallest or largest object, or the one closest to another object). We further show that incorporating chain of thought reasoning allows RT-2 to perform multi-stage semantic reasoning, for example figuring out which object to pick up for use as an improvised hammer (a rock), or which type of drink is best suited for someone who is tired (an energy drink).

Giulia Vezzani

Giulia Vezzani
Staff Research Engineer
Google DeepMind
RoboCat: A self-improving generalist for robotic manipulation
Thu, Nov 23, 2023
Abstract

The ability to leverage heterogeneous robotic experience from different robots and tasks to quickly master novel skills and embodiments has the potential to transform robot learning. Inspired by recent advances in foundation models for vision and language, we propose a multi-embodiment, multi-task generalist agent for robotic manipulation. This agent, named RoboCat, is a visual goal-conditioned decision transformer capable of consuming action-labelled visual experience. This data spans a large repertoire of motor control skills from simulated and real robotic arms with varying sets of observations and actions. With RoboCat, we demonstrate the ability to generalise to new tasks and robots, both zero-shot as well as through adaptation using only 100–1000 examples for the target task. We also show how a trained model itself can be used to generate data for subsequent training iterations, thus providing a basic building block for an autonomous improvement loop. We investigate the agent’s capabilities, with large-scale evaluations both in simulation and on three different real robot embodiments. We find that as we grow and diversify its training data, RoboCat not only shows signs of cross-task transfer, but also becomes more efficient at adapting to new tasks.

Joanne Truong

Sim2Robot: Training Robots for the Real-World with Imperfect Simulators
Thu, Nov 30, 2023
Abstract

The goal of AI is to “construct useful intelligent systems”, such as mobile robots to assist in our day-to-day lives (e.g., home assistant robots tidying the house, or robots delivering packages from one building to another). However, training robots in the real-world can be slow, dangerous, expensive, and difficult to reproduce. Thus, one paradigm in robot learning is to leverage simulation for training robots (where gathering experience is scalable, safe, cheap, and reproducible) before being deployed in the real world. In this talk, I will present methods to leverage imperfect simulators for training robots to act in the real-world (navigation, pick-and-place). First, I will talk about how to use large scale learning in simulation to learn high and low-level robot skills by leveraging hierarchical robot control and abstracted physics. Next, I will talk about how we can combine these robot skills and leverage foundation models to enable open vocabulary mobile pick and place. Finally, I will talk about how we can also generalize these learned navigation skills to out-of-distribution environments zero-shot by leveraging additional context information of the environment.

Ze Yang

Ze Yang
PhD Student
University of Toronto
Senior Research Scientist
Waabi
Learning in-the-wild Sensor Simulation for Autonomous Driving
Thu, Dec 7, 2023
Abstract

Autonomous driving will revolutionize our world and reshape the future of mobility and transportation. The overarching goal of self-driving vehicles (SDVs) is to safely maneuver in diverse environments without human intervention. To accomplish this, we need to rigorously evaluate autonomy systems and train them to handle all possible situations they might encounter in the real world. Sensor simulation emerges to play a crucial role in the realm of autonomous driving. By creating accurate digital replicas of real-world environments and generating simulated sensor data, it facilitates the development and evaluation of robotic systems in a safe, controlled, reactive, and cost-effective manner. In this talk, I will present our recent efforts toward this goal. I will begin with how we construct controllable and realistic digital twins using real-world data. Then I will detail the manipulation of actors, scenes, and environments to create novel scenarios, efficiently rendering them to generate realistic multi-sensor simulations that enhance autonomous systems. Finally, I will delve into our methodology for measuring the simulator's realism within the context of autonomous systems.