Atom-by-atom protein generation and beyond with language models

08/16/2023
by   Daniel Flam-Shepherd, et al.
0

Protein language models learn powerful representations directly from sequences of amino acids. However, they are constrained to generate proteins with only the set of amino acids represented in their vocabulary. In contrast, chemical language models learn atom-level representations of smaller molecules that include every atom, bond, and ring. In this work, we show that chemical language models can learn atom-level representations of proteins enabling protein generation unconstrained to the standard genetic code and far beyond it. In doing so, we show that language models can generate entire proteins atom by atom – effectively learning the multiple hierarchical layers of molecular information that define proteins from their primary sequence to their secondary, and tertiary structure. We demonstrate language models are able to explore beyond protein space – generating proteins with modified sidechains that form unnatural amino acids. Even further, we find that language models can explore chemical space and protein space simultaneously and generate novel examples of protein-drug conjugates. The results demonstrate the potential for biomolecular design at the atom level using language models.

READ FULL TEXT

page 9

page 10

page 15

research
05/09/2023

Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files

Language models are powerful tools for molecular design. Currently, the ...
research
03/29/2023

ProtFIM: Fill-in-Middle Protein Sequence Design via Protein Language Models

Protein language models (pLMs), pre-trained via causal language modeling...
research
09/02/2022

Exploiting Pretrained Biochemical Language Models for Targeted Drug Design

Motivation: The development of novel compounds targeting proteins of int...
research
05/03/2023

Exploring the Protein Sequence Space with Global Generative Models

Recent advancements in specialized large-scale architectures for trainin...
research
11/18/2022

Protein language model rescue mutations highlight variant effects and structure in clinically relevant genes

Despite being self-supervised, protein language models have shown remark...
research
05/25/2023

Explainability Techniques for Chemical Language Models

Explainability techniques are crucial in gaining insights into the reaso...
research
03/14/2020

Lattice protein design using Bayesian learning

A novel protein design method using Bayesian learning is proposed in thi...

Please sign up or login with your details

Forgot password? Click here to reset