DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation

by   Xing Shen, et al.

Binary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28× 28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the training complexity. In this paper, we propose a new mask representation by applying the discrete cosine transform(DCT) to encode the high-resolution binary grid mask into a compact vector. Our method, termed DCT-Mask, could be easily integrated into most pixel-based instance segmentation methods. Without any bells and whistles, DCT-Mask yields significant gains on different frameworks, backbones, datasets, and training schedules. It does not require any pre-processing or pre-training, and almost no harm to the running speed. Especially, for higher-quality annotations and more complex backbones, our method has a greater improvement. Moreover, we analyze the performance of our method from the perspective of the quality of mask representation. The main reason why DCT-Mask works well is that it obtains a high-quality mask representation with low complexity. Code will be made available.


page 3

page 11


BoxInst: High-Performance Instance Segmentation with Box Annotations

We present a high-performance method that can achieve mask-level instanc...

Mask Transfiner for High-Quality Instance Segmentation

Two-stage and query-based instance segmentation methods have achieved re...

Instance Segmentation by Jointly Optimizing Spatial Embeddings and Clustering Bandwidth

Current state-of-the-art instance segmentation methods are not suited fo...

EmbedMask: Embedding Coupling for One-stage Instance Segmentation

Current instance segmentation methods can be categorized into segmentati...

A Histogram Thresholding Improvement to Mask R-CNN for Scalable Segmentation of New and Old Rural Buildings

Mapping new and old buildings are of great significance for understandin...

Video Mask Transfiner for High-Quality Video Instance Segmentation

While Video Instance Segmentation (VIS) has seen rapid progress, current...

SODAR: Segmenting Objects by DynamicallyAggregating Neighboring Mask Representations

Recent state-of-the-art one-stage instance segmentation model SOLO divid...