Speaker

Talk

YouTube link

Florian Shkurti

Robot Videography from Human Specifications
Fri, Sep 4, 2020 · 4:00 PM ET
Abstract

We consider the problem of enabling robots to help with scientific discovery. We want robots that autonomously navigate in unstructured 3D environments, alongside environmental scientists, to help them record footage that they deem scientifically relevant. How can scientists efficiently specify the type of visual data that they want their robots to record? How should robots explore unknown natural environments according to that specification? We address these questions through representation learning methods that enable one-shot informed visual search in unknown environments. Our method can be interpreted as a way to infer the scientist's reward function over visual content. Time permitting, I will also discuss recent progress from our group in terms of continual learning, and its potential to be used in lifelong robot experiments.

Valentin Peretroukhin

Valentin Peretroukhin
Postdoctoral associate
MIT
Representing rotation in deep learning
Fri, Sep 18, 2020 · 4:00 PM ET
Abstract

Estimating rigid-body rotation constitutes one of the core challenges in robot perception. Much recent research has focused on applying the tools of modern deep learning to replace or improve classical algorithms that estimate rotation such as visual and/or inertial odometry, SLAM, and object pose estimation. In order to 'learn' rotations, however, one must deal with the non-trivial (and beautiful) structure of the Special Orthogonal Matrix Lie Group, SO(3). In this talk, I will present my work on using deep neural networks to infer rotations using the machinery of Lie groups and Lie algebra, with a particular emphasis on computing estimates with aleatoric and epistemic uncertainty. Notably, I will outline recent work on a novel representation of rotation based on symmetric matrices that, owing to its smooth structure and intimate connection to the Bingham distribution, is particularly suited to deep rotation learning.

Ankur Handa

Ankur Handa
Research scientist
NVIDIA
DexPilot - Vision-based teleoperation of dextrous robotic hand-arm system
Fri, Sep 25, 2020 · 4:00 PM ET
Abstract

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, current teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usually provide reduced degrees of control. Herein, a low-cost, vision-based teleoperation system, DexPilot, was developed that allows for complete control over the full 23 DoA robotic system by merely observing the bare human hand. DexPilot enables operators to carry out a variety of complex manipulation tasks that go beyond simple pick-and-place operations. This allows for the collection of high dimensional, multi-modality, state-action data that can be leveraged in the future to learn sensorimotor policies for challenging manipulation tasks. The system performance was measured through speed and reliability metrics across two human demonstrators on a variety of tasks.

Shubham Tulsiani

Self-supervised interaction with the physical world
Fri, Oct 9, 2020 · 4:00 PM ET
Abstract

We live in a physical world, and any artificially intelligent agents must understand and act in it. In this talk, I will present some recent efforts addressing two central questions relevant to the pursuit of such intelligent agents: a) what representations should perception systems infer to guide action?, and b) what actions can be informative for learning better perception?

I will argue that building systems that (implicitly or explicitly) infer the physical structure underlying visual percepts can allow acting efficiently, and that incorporating our knowledge about the laws of physics can help bypass the need for tedious manual supervision for learning. I will then discuss how such agents can seek informative exploratory actions by leveraging the inherent multi-modality in the perceptual input.

Ronald Clark

Representation Learning for 3D Vision
Fri, Oct 16, 2020 · 4:00 PM ET
Abstract

For robots to understand and interact with the world around them, they need access to a good representation of their environment. In this talk we’ll look at how deep learning can help to build such a representation. We’ll look at how generative priors can be used to improve the representation learning capabilities of variational autoencoders (VAEs). We’ll also investigate how 3D geometry can help build reliable spatial memories. Finally, we’ll see how these representations can be used for accurate state estimation. In essence, representation learning for 3D vision opens up many exciting possibilities and I hope this talk will spark some interesting ideas in this direction.

Lerrel Pinto

Lerrel Pinto
Assistant professor
New York University
Robot learning in the wild
Fri, Oct 23, 2020 · 3:00 PM ET
Abstract

While robotics has made tremendous progress over the last few decades, most success stories are still limited to carefully engineered and precisely modeled environments. Interestingly, one of the most significant successes in the last decade of AI has been the use of Machine Learning (ML) to generalize and robustly handle diverse situations. So why don't we just apply current learning algorithms to robots? The biggest reason is a complicated relationship between data and robotics. In other fields of AI such as computer vision, we were able to collect diverse real-world, large-scale data with lots of supervision. These three key ingredients which fueled the success of deep learning in other fields are the key bottlenecks in robotics. We do not have millions of training examples in robots; it is unclear how to supervise robots and most importantly, simulation/lab data is not real-world and diverse. My research has focused on rethinking the relationship between data and robotics to fuel the success of robot learning. Specifically, in this talk, I will discuss three aspects of data that will bring us closer to generalizable robotics: (a) size of data we can collect, (b) amount of supervisory signal we can extract, and (c) diversity of data we can get from robots.

Mustafa Mukadam

Differentiable motion planning
Fri, Nov 6, 2020 · 4:00 PM ET
Abstract

In robotics and more specifically motion planning, recent debates have transitioned from a binary choice between hand crafted priors vs deep learning, towards how to leverage both ends of this spectrum. Best ways to combine them and strike a balance remains an open research question. In this talk, I will present an inference based approach to motion planning built using factor graphs and show how this setup can be used as a solid and flexible foundation in exploring the role of learning in adding value over traditional methods. We will arrive at a fully differentiable approach that can be trained end-to-end while incorporating prior knowledge.

Shuran Song

Shuran Song
Assistant professor
Columbia University
Active scene understanding with robot interactions
Fri, Nov 20, 2020 · 4:00 PM ET
Abstract

Most computer vision algorithms are built with the goal to understand the physical world. Yet, as reflected in standard vision benchmarks and datasets, these algorithms continue to assume the role of a passive observer -- only watching static images or videos, without the ability to interact with the environment. This assumption becomes a fundamental limitation for applications in robotics, where systems are intrinsically built to actively engage with the physical world. In this talk, I will present some recent work from my group that demonstrates how we can enable robots to leverage their ability to interact with the environment in order to better understand what they see: from discovering objects' identity and 3D geometry to learning their physical properties. We will demonstrate how the learned knowledge can be used to facilitate downstream manipulation tasks. Finally, I will discuss a few open research directions in the area of active scene understanding.

Angela Schoellig

Combining models and data for increased performance and safety in robotics
Fri, Dec 4, 2020 · 4:10 PM ET
Abstract

The ultimate promise of robotics is to design devices that can physically interact with the world. To date, robots have been primarily deployed in highly structured and predictable environments. However, we envision the next generation of robots (ranging from self-driving and -flying vehicles to robot assistants) to operate in unpredictable and generally unknown environments and alongside humans. This challenges current robot algorithms, which have been largely based on a-priori knowledge about the system and its environment. While research has shown that robots are able to learn new skills from experience and adapt to unknown situations, these results have been mostly limited to learning single tasks, and demonstrated in simulation or structured lab settings. The next challenge is to enable robot learning in real-world application scenarios. This will require versatile, data-efficient and online learning algorithms that guarantee safety. It will also require to answer the fundamental question of how to design learning architectures for dynamic and interactive agents. This talk will highlight our recent progress in combining learning methods with formal results from control theory. By combining models with data, our algorithms achieve adaptation to changing conditions during long-term operation, data-efficient multi-robot, multi-task transfer learning, and safe reinforcement learning. We demonstrate our algorithms in vision-based off-road driving and drone flight experiments, as well as on mobile manipulators.