SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition

01/08/2020
by   Zhixuan Lin, et al.
28

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either based on spatial-attention or scene-mixture approaches and limited in scalability which is a main obstacle towards modeling real-world scenes. In this paper, we propose a generative latent variable model, called SPACE, that provides a unified probabilistic modeling framework that combines the best of spatial-attention and scene-mixture approaches. SPACE can explicitly provide factorized object representations for foreground objects while also decomposing background segments of complex morphology. Previous models are good at either of these, but not both. SPACE also resolves the scalability problems of previous methods by incorporating parallel spatial-attention and thus is applicable to scenes with a large number of objects without performance degradations. We show through experiments on Atari and 3D-Rooms that SPACE achieves the above properties consistently in comparison to SPAIR, IODINE, and GENESIS. Results of our experiments can be found on our project website: https://sites.google.com/view/space-project-page

READ FULL TEXT

page 6

page 7

page 12

page 13

page 14

research
10/06/2019

Scalable Object-Oriented Sequential Generative Models

The main limitation of previous approaches to unsupervised sequential ob...
research
01/22/2019

MONet: Unsupervised Scene Decomposition and Representation

The ability to decompose scenes in terms of abstract building blocks is ...
research
06/03/2021

GMAIR: Unsupervised Object Detection Based on Spatial Attention and Gaussian Mixture

Recent studies on unsupervised object detection based on spatial attenti...
research
06/11/2020

Learning to Infer 3D Object Models from Images

A crucial ability of human intelligence is to build up models of individ...
research
04/19/2023

Learning Temporal Distribution and Spatial Correlation for Universal Moving Object Segmentation

Universal moving object segmentation aims to provide a general model for...
research
06/27/2012

Learning Object Arrangements in 3D Scenes using Human Context

We consider the problem of learning object arrangements in a 3D scene. T...
research
04/30/2023

Object-Centric Voxelization of Dynamic Scenes via Inverse Neural Rendering

Understanding the compositional dynamics of the world in unsupervised 3D...

Please sign up or login with your details

Forgot password? Click here to reset