Positional Encoding as Spatial Inductive Bias in GANs

12/09/2020
by   Rui Xu, et al.
0

SinGAN shows impressive capability in learning internal patch distribution despite its limited effective receptive field. We are interested in knowing how such a translation-invariant convolutional generator could capture the global structure with just a spatially i.i.d. input. In this work, taking SinGAN and StyleGAN2 as examples, we show that such capability, to a large extent, is brought by the implicit positional encoding when using zero padding in the generators. Such positional encoding is indispensable for generating images with high fidelity. The same phenomenon is observed in other generative architectures such as DCGAN and PGGAN. We further show that zero padding leads to an unbalanced spatial bias with a vague relation between locations. To offer a better spatial inductive bias, we investigate alternative positional encodings and analyze their effects. Based on a more flexible positional encoding explicitly, we propose a new multi-scale training strategy and demonstrate its effectiveness in the state-of-the-art unconditional generator StyleGAN2. Besides, the explicit spatial inductive bias substantially improve SinGAN for more versatile image manipulation.

READ FULL TEXT

page 8

page 17

page 18

page 19

page 20

page 21

page 22

page 23

research
08/03/2021

Toward Spatially Unbiased Generative Models

Recent image generation models show remarkable generation performance. H...
research
03/16/2020

On Translation Invariance in CNNs: Convolutional Layers can Exploit Absolute Spatial Location

In this paper we challenge the common assumption that convolutional laye...
research
12/05/2020

Spatially-Adaptive Pixelwise Networks for Fast Image Translation

We introduce a new generator architecture, aimed at fast and efficient h...
research
08/18/2022

The 8-Point Algorithm as an Inductive Bias for Relative Pose Prediction by ViTs

We present a simple baseline for directly estimating the relative pose (...
research
07/17/2023

Cumulative Spatial Knowledge Distillation for Vision Transformers

Distilling knowledge from convolutional neural networks (CNNs) is a doub...
research
01/18/2019

Learning Spatial Pyramid Attentive Pooling in Image Synthesis and Image-to-Image Translation

Image synthesis and image-to-image translation are two important generat...
research
10/29/2019

Convolutional Conditional Neural Processes

We introduce the Convolutional Conditional Neural Process (ConvCNP), a n...

Please sign up or login with your details

Forgot password? Click here to reset