GEOM: Energy-annotated molecular conformations for property prediction and molecular generation

06/09/2020
by   Simon Axelrod, et al.
0

Machine learning outperforms traditional approaches in many molecular design tasks. Although molecules are often thought of as 2D graphs, they in fact consist of an ensemble of inter-converting 3D structures called conformers. Molecular properties arise from the contribution of many conformers, and in the case of a drug binding a target, may be due mainly to a few distinct members. Molecular representations in machine learning are typically based on either one single 3D conformer or on a 2D graph that strips geometrical information. No reference datasets exist that connect these graph and point cloud ensemble representations. Here, we use first-principles simulations to annotate over 400,000 molecules with the ensemble of geometries they span. The Geometrical Embedding Of Molecules (GEOM) dataset contains over 33 million molecular conformers labeled with their relative energies and statistical probabilities at room temperature. This dataset will assist benchmarking and transfer learning in two classes of tasks: inferring 3D properties from 2D molecular graphs, and developing generative models to sample 3D conformations.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/07/2019

Transfer Learning Using Ensemble Neural Nets for Organic Solar Cell Screening

Organic Solar Cells are a promising technology for solving the clean ene...
research
06/22/2019

Alchemy: A Quantum Chemistry Dataset for Benchmarking AI Models

We introduce a new molecular dataset, named Alchemy, for developing mach...
research
02/01/2022

Scalable Fragment-Based 3D Molecular Design with Reinforcement Learning

Machine learning has the potential to automate molecular design and dras...
research
05/06/2022

Transferring Chemical and Energetic Knowledge Between Molecular Systems with Machine Learning

Predicting structural and energetic properties of a molecular system is ...
research
06/13/2023

Von Mises Mixture Distributions for Molecular Conformation Generation

Molecules are frequently represented as graphs, but the underlying 3D mo...
research
08/09/2020

Augmenting Molecular Images with Vector Representations as a Featurization Technique for Drug Classification

One of the key steps in building deep learning systems for drug classifi...
research
02/11/2020

Predicting drug properties with parameter-free machine learning: Pareto-Optimal Embedded Modeling (POEM)

The prediction of absorption, distribution, metabolism, excretion, and t...

Please sign up or login with your details

Forgot password? Click here to reset