SARS-CoV-2 Coronavirus Data Compression Benchmark

12/21/2020
by   Innar Liiv, et al.
0

This paper introduces a lossless data compression competition that benchmarks solutions (computer programs) by the compressed size of the 44,981 concatenated SARS-CoV-2 sequences, with a total uncompressed size of 1,339,868,341 bytes. The data, downloaded on 13 December 2020, from the severe acute respiratory syndrome coronavirus 2 data hub of ncbi.nlm.nih.gov is presented in FASTA and 2Bit format. The aim of this competition is to encourage multidisciplinary research to find the shortest lossless description for the sequences and to demonstrate that data compression can serve as an objective and repeatable measure to align scientific breakthroughs across disciplines. The shortest description of the data is the best model; therefore, further reducing the size of this description requires a fundamental understanding of the underlying context and data. This paper presents preliminary results with multiple well-known compression algorithms for baseline measurements, and insights regarding promising research avenues. The competition's progress will be reported at <https://coronavirus.innar.com>, and the benchmark is open for all to participate and contribute.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/09/2023

Hierarchical Autoencoder-based Lossy Compression for Large-scale High-resolution Scientific Data

Lossy compression has become an important technique to reduce data size ...
research
07/13/2020

Local Editing in LZ-End Compressed Data

This paper presents an algorithm for the modification of data compressed...
research
08/25/2022

LightAMR format standard and lossless compression algorithms for adaptive mesh refinement grids: RAMSES use case

The evolution of parallel I/O library as well as new concepts such as 'i...
research
06/11/2018

Compression of phase-only holograms with JPEG standard and deep learning

It is a critical issue to reduce the enormous amount of data in the proc...
research
03/29/2021

Measuring Sample Efficiency and Generalization in Reinforcement Learning Benchmarks: NeurIPS 2020 Procgen Benchmark

The NeurIPS 2020 Procgen Competition was designed as a centralized bench...
research
08/16/2011

A Machine Learning Perspective on Predictive Coding with PAQ

PAQ8 is an open source lossless data compression algorithm that currentl...
research
11/09/2021

An Examination of Sport Climbing's Competition Format and Scoring System

Sport climbing, which made its Olympic debut at the 2020 Summer Games, g...

Please sign up or login with your details

Forgot password? Click here to reset