FlowFormer: A Transformer Architecture and Its Masked Cost Volume Autoencoding for Optical Flow

06/08/2023
by   Zhaoyang Huang, et al.
0

This paper introduces a novel transformer-based network architecture, FlowFormer, along with the Masked Cost Volume AutoEncoding (MCVA) for pretraining it to tackle the problem of optical flow estimation. FlowFormer tokenizes the 4D cost-volume built from the source-target image pair and iteratively refines flow estimation with a cost-volume encoder-decoder architecture. The cost-volume encoder derives a cost memory with alternate-group transformer (AGT) layers in a latent space and the decoder recurrently decodes flow from the cost memory with dynamic positional cost queries. On the Sintel benchmark, FlowFormer architecture achieves 1.16 and 2.09 average end-point-error (AEPE) on the clean and final pass, a 16.5% and 15.5% error reduction from the GMA (1.388 and 2.47). MCVA enhances FlowFormer by pretraining the cost-volume encoder with a masked autoencoding scheme, which further unleashes the capability of FlowFormer with unlabeled data. This is especially critical in optical flow estimation because ground truth flows are more expensive to acquire than labels in other vision tasks. MCVA improves FlowFormer all-sided and FlowFormer+MCVA ranks 1st among all published methods on both Sintel and KITTI-2015 benchmarks and achieves the best generalization performance. Specifically, FlowFormer+MCVA achieves 1.07 and 1.94 AEPE on the Sintel benchmark, leading to 7.76% and 7.18% error reductions from FlowFormer.

READ FULL TEXT

page 2

page 5

page 7

page 11

page 16

page 17

research
03/30/2022

FlowFormer: A Transformer Architecture for Optical Flow

We introduce Optical Flow TransFormer (FlowFormer), a transformer-based ...
research
03/02/2023

FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation

FlowFormer introduces a transformer architecture into optical flow estim...
research
04/17/2023

LLA-FLOW: A Lightweight Local Aggregation on Cost Volume for Optical Flow Estimation

Lack of texture often causes ambiguity in matching, and handling this is...
research
06/09/2023

DIFT: Dynamic Iterative Field Transforms for Memory Efficient Optical Flow

Recent advancements in neural network-based optical flow estimation ofte...
research
02/20/2018

Devon: Deformable Volume Network for Learning Optical Flow

We propose a lightweight neural network model, Deformable Volume Network...
research
12/20/2022

CGCV:Context Guided Correlation Volume for Optical Flow Neural Networks

Optical flow, which computes the apparent motion from a pair of video fr...
research
04/12/2023

MED-VT: Multiscale Encoder-Decoder Video Transformer with Application to Object Segmentation

Multiscale video transformers have been explored in a wide variety of vi...

Please sign up or login with your details

Forgot password? Click here to reset