AMPL: A Data-Driven Modeling Pipeline for Drug Discovery

11/13/2019
by   Amanda J. Minnich, et al.
51

One of the key requirements for incorporating machine learning into the drug discovery process is complete reproducibility and traceability of the model building and evaluation process. With this in mind, we have developed an end-to-end modular and extensible software pipeline for building and sharing machine learning models that predict key pharma-relevant parameters. The ATOM Modeling PipeLine, or AMPL, extends the functionality of the open source library DeepChem and supports an array of machine learning and molecular featurization tools. We have benchmarked AMPL on a large collection of pharmaceutical datasets covering a wide range of parameters. Our key findings include: Physicochemical descriptors and deep learning-based graph representations are significantly better than traditional fingerprints to characterize molecular features Dataset size is directly correlated to performance of prediction: single-task deep learning models only outperform shallow learners if there is enough data. Likewise, data set size has a direct impact of model predictivity independently of comprehensive hyperparameter model tuning. Our findings point to the need for public dataset integration or multi-task/transfer learning approaches. DeepChem uncertainty quantification (UQ) analysis may help identify model error; however, efficacy of UQ to filter predictions varies considerably between datasets and model types. This software is open source and available for download at http://github.com/ATOMconsortium/AMPL.

READ FULL TEXT

page 4

page 12

page 14

page 16

page 18

research
12/09/2020

Utilising Graph Machine Learning within Drug Discovery and Development

Graph Machine Learning (GML) is receiving growing interest within the ph...
research
04/30/2020

A Systematic Approach to Featurization for Cancer Drug Sensitivity Predictions with Deep Learning

By combining various cancer cell line (CCL) drug screening panels, the s...
research
02/14/2023

Do Deep Learning Methods Really Perform Better in Molecular Conformation Generation?

Molecular conformation generation (MCG) is a fundamental and important p...
research
12/10/2019

libmolgrid: GPU Accelerated Molecular Gridding for Deep Learning Applications

There are many ways to represent a molecule as input to a machine learni...
research
01/22/2021

Nigraha: Machine-learning based pipeline to identify and evaluate planet candidates from TESS

The Transiting Exoplanet Survey Satellite (TESS) has now been operationa...
research
03/24/2023

Machine Guided Discovery of Novel Carbon Capture Solvents

The increasing importance of carbon capture technologies for deployment ...
research
12/27/2021

ToxTree: descriptor-based machine learning models for both hERG and Nav1.5 cardiotoxicity liability predictions

Drug-mediated blockade of the voltage-gated potassium channel(hERG) and ...

Please sign up or login with your details

Forgot password? Click here to reset