3D Neural Scene Representations for Visuomotor Control

07/08/2021
by   Yunzhu Li, et al.
10

Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us to perform a wide range of manipulation tasks that are far beyond the reach of current robots. In this work, we desire to learn models for dynamic 3D scenes purely from 2D visual observations. Our model combines Neural Radiance Fields (NeRF) and time contrastive learning with an autoencoding framework, which learns viewpoint-invariant 3D-aware scene representations. We show that a dynamics model, constructed over the learned representation space, enables visuomotor control for challenging manipulation tasks involving both rigid bodies and fluids, where the target is specified in a viewpoint different from what the robot operates on. When coupled with an auto-decoding framework, it can even support goal specification from camera viewpoints that are outside the training distribution. We further demonstrate the richness of the learned 3D dynamics model by performing future prediction and novel view synthesis. Finally, we provide detailed ablation studies regarding different system designs and qualitative analysis of the learned representations.

READ FULL TEXT

page 2

page 3

page 6

page 7

page 8

page 14

page 15

research
10/03/2018

Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids

Real-life control tasks involve matter of various substances---rigid or ...
research
04/22/2023

3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes

Given a visual scene, humans have strong intuitions about how a scene ca...
research
12/12/2022

MIRA: Mental Imagery for Robotic Affordances

Humans form mental images of 3D scenes to support counterfactual imagina...
research
11/12/2020

3D-OES: Viewpoint-Invariant Object-Factorized Environment Simulators

We propose an action-conditioned dynamics model that predicts scene chan...
research
11/03/2020

Learning 3D Dynamic Scene Representations for Robot Manipulation

3D scene representation for robot manipulation should capture three key ...
research
04/14/2022

Manually Acquiring Targets from Multiple Viewpoints Using Video Feedback

Objective: The effect of camera viewpoint was studied when performing vi...
research
09/07/2018

Neural Allocentric Intuitive Physics Prediction from Real Videos

Humans are able to make rich predictions about the future dynamics of ph...

Please sign up or login with your details

Forgot password? Click here to reset