Hallucinating Pose-Compatible Scenes

12/13/2021
by   Tim Brooks, et al.
3

What does human pose tell us about a scene? We propose a task to answer this question: given human pose as input, hallucinate a compatible scene. Subtle cues captured by human pose – action semantics, environment affordances, object interactions – provide surprising insight into which scenes are compatible. We present a large-scale generative adversarial network for pose-conditioned scene generation. We significantly scale the size and complexity of training data, curating a massive meta-dataset containing over 19 million frames of humans in everyday environments. We double the capacity of our model with respect to StyleGAN2 to handle such complex data, and design a pose conditioning mechanism that drives our model to learn the nuanced relationship between pose and scene. We leverage our trained model for various applications: hallucinating pose-compatible scene(s) with or without humans, visualizing incompatible scenes and poses, placing a person from one generated image into another scene, and animating pose. Our model produces diverse samples and outperforms pose-conditioned StyleGAN2 and Pix2Pix baselines in terms of accurate human placement (percent of correct keypoints) and image quality (Frechet inception distance).

READ FULL TEXT

page 4

page 5

page 6

page 14

page 15

page 16

page 17

page 18

research
05/01/2020

Adversarial Synthesis of Human Pose from Text

This work introduces the novel task of human pose synthesis from text. I...
research
12/05/2019

Generating 3D People in Scenes without People

We present a fully-automatic system that takes a 3D scene and generates ...
research
04/09/2018

Binge Watching: Scaling Affordance Learning from Sitcoms

In recent years, there has been a renewed interest in jointly modeling p...
research
12/23/2021

HSPACE: Synthetic Parametric Humans Animated in Complex Environments

Advances in the state of the art for 3d human sensing are currently limi...
research
08/04/2023

Scene-aware Human Pose Generation using Transformer

Affordance learning considers the interaction opportunities for an actor...
research
12/01/2021

PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound

Reconstructing the 3D pose of a person in metric scale from a single vie...
research
07/28/2022

The One Where They Reconstructed 3D Humans and Environments in TV Shows

TV shows depict a wide variety of human behaviors and have been studied ...

Please sign up or login with your details

Forgot password? Click here to reset