ProteinNet: a standardized data set for machine learning of protein structure

02/01/2019
by   Mohammed AlQuraishi, et al.
0

Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been driven by standardized data sets that facilitate fair assessment of new methods and lower the barrier to entry for non-domain experts. While data sets of protein sequence and structure exist, they lack certain components critical for machine learning, including high-quality multiple sequence alignments and insulated training / validation splits that account for deep but only weakly detectable homology across protein space. We have created the ProteinNet series of data sets to provide a standardized mechanism for training and assessing data-driven models of protein sequence-structure relationships. ProteinNet integrates sequence, structure, and evolutionary information in programmatically accessible file formats tailored for machine learning frameworks. Multiple sequence alignments of all structurally characterized proteins were created using substantial high-performance computing resources. Standardized data splits were also generated to emulate the difficulty of past CASP (Critical Assessment of protein Structure Prediction) experiments by resetting protein sequence and structure space to the historical states that preceded six prior CASPs. Utilizing sensitive evolution-based distance metrics to segregate distantly related proteins, we have additionally created validation sets distinct from the official CASP sets that faithfully mimic their difficulty. ProteinNet thus represents a comprehensive and accessible resource for training and assessing machine-learned models of protein structure.

READ FULL TEXT
POST COMMENT

Comments

There are no comments yet.

Authors

page 1

page 6

07/16/2020

Deep Learning in Protein Structural Modeling and Design

Deep learning is catalyzing a scientific revolution fueled by big data, ...
03/23/2022

A Supervised Machine Learning Approach for Sequence Based Protein-protein Interaction (PPI) Prediction

Computational protein-protein interaction (PPI) prediction techniques ca...
10/16/2020

SidechainNet: An All-Atom Protein Structure Dataset for Machine Learning

Despite recent advancements in deep learning methods for protein structu...
11/09/2019

Accurate Protein Structure Prediction by Embeddings and Deep Learning Representations

Proteins are the major building blocks of life, and actuators of almost ...
09/18/2015

Evaluation of Protein-protein Interaction Predictors with Noisy Partially Labeled Data Sets

Protein-protein interaction (PPI) prediction is an important problem in ...
11/17/2018

High Quality Prediction of Protein Q8 Secondary Structure by Diverse Neural Network Architectures

We tackle the problem of protein secondary structure prediction using a ...
06/17/2018

MCP: a multi-component learning machine for prediction of protein secondary structure

Proteins biological function is tightly connected to its specific 3D str...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.