Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos

05/25/2023
by   Matthew Chang, et al.
0

The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents a nuisance. However, often hands also provide a valuable signal, e.g. the hand pose may suggest what kind of object is being held. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving inpainting quality on egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos.

READ FULL TEXT

page 2

page 3

page 6

page 9

page 18

page 19

research
05/04/2023

Learning Hand-Held Object Reconstruction from In-The-Wild Videos

Prior works for reconstructing hand-held objects from a single image rel...
research
02/01/2022

DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video

Dexterous multi-fingered robotic hands have a formidable action space, y...
research
08/15/2021

Occlusion-Aware Video Object Inpainting

Conventional video inpainting is neither object-oriented nor occlusion-a...
research
04/23/2021

H2O: A Benchmark for Visual Human-human Object Handover Analysis

Object handover is a common human collaboration behavior that attracts a...
research
08/07/2022

Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications

Egocentric videos offer fine-grained information for high-fidelity model...
research
12/22/2014

Occlusion Edge Detection in RGB-D Frames using Deep Convolutional Networks

Occlusion edges in images which correspond to range discontinuity in the...
research
08/18/2023

O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model

Occlusion is a common issue in 3D reconstruction from RGB-D videos, ofte...

Please sign up or login with your details

Forgot password? Click here to reset