IP Library › Granted Patent US 12,573,475
Granted Patent B2
US 12,573,475 · App. 18/199,144 · Granted Mar 10, 2026

Protein sequence and structure generation with denoising diffusion probabilistic models

Inventors: Namrata Anand (Menlo Park, CA); Tudor Achim (Menlo Park, CA)
Assignee: Diffuse Bio, Inc.
G16B40/00G16B15/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,475
App. No.
18/199,144
Granted
Mar 10, 2026
Kind
B2
Abstract

Training a protein diffusion model includes receiving a representation of a protein as training data, the representation comprising at least three dimensions. It further includes training a protein diffusion model at least in part by performing rotational diffusion based at least in part on the representation of the protein. Generating proteins includes receiving protein conditioning information. It further includes, based at least in part on the protein conditioning information, performing conditional sampling of a protein diffusion model. The protein diffusion model is trained at least in part by performing rotational diffusion. Based at least in part on the conditional sampling of the protein diffusion model, the protein diffusion model generates one or more of a protein structure or a protein sequence.

Claims (27)

1 . A method, comprising:

expressing, in a laboratory setting, an amino acid sequence, wherein the amino acid sequence was generated as output of a protein diffusion model, wherein the amino acid sequence is predicted by the protein diffusion model to fold to a desired backbone structure, and wherein generating the amino acid sequence using the protein diffusion model comprises:

receiving a representation of a protein as training data, the representation comprising at least three dimensions, and wherein the representation of the protein comprises, for an atom of a backbone structure of the protein, a corresponding coordinate and a corresponding coordinate frame;

training the protein diffusion model on at least one graphics processing unit (GPU) at least in part by performing rotational diffusion based at least in part on the representation of the protein; and

using the protein diffusion model to generate a binder to a target protein included in protein conditioning information that is received as input, wherein an evaluation of the binder generated for the target protein included in the protein conditioning information is performed, including evaluating a delta energy corresponding to interacting of the binder generated using the protein diffusion model and the target protein included in the protein conditioning information received as input, and wherein generating the binder comprises generating the amino acid sequence.

2 . The method of claim 1 , wherein performing the rotational diffusion comprises performing interpolation between rotations.

3 . The method of claim 2 , wherein performing the interpolation comprises interpolating between rotational frames of reference.

4 . The method of claim 1 , wherein training the protein diffusion model comprises determining one or more parameters of the protein diffusion model based at least in part on the rotational diffusion and computing of a loss.

5 . The method of claim 4 , wherein computing the loss comprises aligning coordinate frames.

6 . The method of claim 1 , wherein the representation of the protein comprises angles associated with rotamers.

7 . A method, comprising:

expressing, in a laboratory setting, an amino acid sequence, wherein the amino acid sequence was generated as output of a protein diffusion model, wherein the amino acid sequence is predicted by the protein diffusion model to fold to a desired backbone structure, and wherein generating the amino acid sequence using the protein diffusion model comprises:

receiving, as input, protein conditioning information including a target protein;

based at least in part on the protein conditioning information, performing conditional sampling of the protein diffusion model, wherein the protein diffusion model is trained on at least one graphics processing unit (GPU) at least in part by performing rotational diffusion based at least in part on a representation of a protein comprising, for an atom of a backbone structure of the protein, a corresponding coordinate and a corresponding coordinate frame;

wherein based at least in part on the conditional sampling of the protein diffusion model, the protein diffusion model generates a binder to the target protein included in the protein conditioning information that is received as input, and wherein generating the binder using the protein diffusion model comprises generating the amino acid sequence; and

performing an evaluation of the binder generated for the target protein included in the protein conditioning information, including evaluating a delta energy corresponding to interacting of the binder generated using the protein diffusion model and the target protein included in the protein conditioning information received as input.

8 . The method of claim 7 , wherein the protein conditioning information comprises an indication of protein residues.

9 . The method of claim 8 , wherein the protein conditioning information comprises an indication of blocks into which the protein residues are divided.

10 . The method of claim 9 , wherein the protein conditioning information comprises a block secondary structure assignment.

11 . The method of claim 9 , wherein the protein conditioning information comprises block adjacency information.

12 . The method of claim 11 , wherein the block adjacency information comprises an indication of adjacency or non-adjacency between two blocks.

13 . The method of claim 10 wherein the protein conditioning information comprises an indication of whether a beta sheet pairing is parallel or anti-parallel.

14 . The method of claim 7 , further comprising:

performing the conditional sampling of the protein diffusion model at least in part by:

determining a plurality of noise samples; and

for each noise sample in the plurality of noise samples, conditionally sampling the protein diffusion model using the noise sample and the protein conditioning information; and

wherein a plurality of binders are generated by the protein diffusion model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2025
From: ANAND, NAMRATA; ACHIM, TUDOR
To: DIFFUSE BIO, INC.
Reel/Frame 072272/0767 →
Continuity (2)
Provisional Application 63343789 · May 19, 2022
Related Publication 20230377690A1 · Nov 23, 2023
References Cited (37)
Torrisi et al. “Deep Learning Methods in Protein Structure Prediction” Computational and Structural Biotechnology Journal (2020) vol. 18, pp. 1301-1310 (Year: 2020). [cited by examiner]
Akpinaroglu et al., Improved antibody structure prediction by deep learning of side chain conformations, BioRxiv, Sep. 17, 2021, pp. 1-16, https://doi.org/10.1101/2021.09.22.461349. [cited by applicant]
Alford et al., The Rosetta all-atom energy function for macromolecular modeling and design, Journal of Chemical Theory and Computation, 13(6), Feb. 7, 2017, pp. 1-35, https://doi.org/10.1101/106054. [cited by applicant]
Anand et al., Generative Modeling for Protein Structures, Advances in Neural Information Processing Systems, 31, 2018, pp. 1-12. [cited by applicant]
Anand et al., Protein sequence design with a learned potential, Nature Communications, 13(1):746, Feb. 8, 2022, pp. 1-11, https://doi.org/10.1038/s41467-022-28313-9. [cited by applicant]
Austin et al., Structured Denoising Diffusion Models in Discrete State-Spaces, Advances in Neural Information Processing Systems, 34, Dec. 6, 2021, pp. 1-33, arXiv:2017.03006v3 [cs.LG]. [cited by applicant]
Berman et al., The Protein Data Bank, Nucleic Acids Research, vol. 28, Issue 1, Jan. 1, 2000, pp. 235-242, https://doi.org/10.1093/nar/28.1.235. [cited by applicant]
Brown et al., Language Models are Few-Shot Learners, Advances in Neural Information Processing Systems, vol. 33, Jul. 22, 2020, pp. 1877-1901, arXiv:2005.14165v4 [cs.CL]. [cited by applicant]
Castro et al., ReLSO: A Transformer-based Model for Latent Space Optimization and Generation of Proteins, arXiv preprint, Jan. 24, 2022, pp. 1-25, arXiv:2201.09948v2 [cs.LG]. [cited by applicant]
Dawson et al., CATH: an expanded resource to predict protein function through structure and sequence, Nucleic Acids Research, vol. 45, Issue D1, Jan. 2017, pp. D289-D295, https://doi.org/10.1093/nar/gkw1098. [cited by applicant]
De Bortoli et al., Riemannian Score-Based Generative Modeling, Advances in Neural Information Processing Systems, Nov. 22, 2022, 50 pages, arXiv:2202.02763 [cs.LG]. [cited by applicant]
Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Proceedings of NAACL-HLT 2019, May 24, 2019, pp. 4171-4186. [cited by applicant]
Dhariwal et al., Diffusion Models Beat GANS on Image Synthesis, Advances in Neural Information Processing Systems, 34, Dec. 6, 2021, pp. 1-15. [cited by applicant]
Du et al., Energy-Based Models for Atomic-Resolution Protein Conformations, ICLR 2020 Conference Paper, Apr. 27, 2020, pp. 1-16, arXiv:2004.13167v1 [cs.LG]. [cited by applicant]
Eguchi et al., IG-VAE: Generative Modeling of Immunoglobulin Proteins by Direct 3D Coordinate Generation, bioRxiv, Jan. 1, 2020, pp. 1-13, https://doi.org/10.1101/2020.08.07.242347. [cited by applicant]
Ferruz et al., A deep unsupervised language model for protein design, bioRxiv, Mar. 12, 2022, pp. 1-14, https://doi.org/10.1101/2022.03.09.483666. [cited by applicant]
Ferruz et al., Towards Controllable Protein Design with Conditional Transformers, arXiv preprint, 2022, pp. 1-17, arXiv:2201.07338. [cited by applicant]
Gao et al., AlphaDesign: A graph protein design method and benchmark on AlphaFoldDB, arXiv preprint, Feb. 12, 2022, 11 pages, arXiv:2202.01079v2 [q-bio.QM]. [cited by applicant]
Ho et al., Denoising Diffusion Probabilistic Models, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Dec. 16, 2020, pp. 1-25, arXiv:2006.11239v2 [cs.LG]. [cited by applicant]
Hsu et al., Learning inverse folding from millions of predicted structures, vioRxiv preprint, Systems Biology, Apr. 10, 2022, pp. 1-22, https://doi.org/10.1101/2022.04.10.487779. [cited by applicant]
Ingraham et al., Generative models for graph-based protein design, Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 1-12. [cited by applicant]
Jing et al., Torsional Diffusion for Molecular Conformer Generation, ICLR2022 Machine Learning for Drug Discovery, Dec. 6, 2022, pp. 1-28, arXiv:2206.01729v2 [physics.chem-ph]. [cited by applicant]
Jumper et al., Highly accurate protein structure prediction with AlphaFold, Nature, vol. 596, Aug. 26, 2021, pp. 583-589. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization”, arXiv preprint, Dec. 22, 2014, pp. 1-15, arXiv:1412.6980v9 [cs.LG]. [cited by applicant]
Kuhlman et al., Native protein sequences are close to optimal for their structures, Proceedings of the National Academy of Sciences, vol. 97, No. 19, Sep. 12, 2000, pp. 10383-10388. [cited by applicant]
Leaver-Fay et al., Scientific Benchmarks for Guiding Macromolecular Energy Function Improvement, Methods in Enzymology, vol. 523, Jan. 1, 2013, pp. 109-143. [cited by applicant]
Lewis et al., Gene3D: Extensive prediction of globular domains in proteins, Nucleic Acids Research, vol. 46, Database Issue, Nov. 3, 2017, pp. D435-D439. [cited by applicant]
Liu et al., Prediction of amino acid side chain conformation using a deep neural network, arXiv preprint, Jul. 26, 2017, 39 pages, arXiv:1707.08381. [cited by applicant]
Loshchilov et al., SGDR: Stochastic Gradient Descent with Warm Restarts, arXiv preprint, Aug. 13, 2016, pp. 1-16, arXiv:1608.03983v5 [CS.LG]. [cited by applicant]
Madani et al., ProGen: Language Modeling for Protein Generation, arXiv preprint, Mar. 8, 2020, 17 pages, arXiv:2004.03497v1 [q-bio.BM]. [cited by applicant]
Mcpartlon et al., An end-to-end deep learning method for rotamer-free protein side-chain packing, bioRxiv preprint, Mar. 14, 2022, pp. 1-29, https://doi.org/10.1101/2022.03.11.483812. [cited by applicant]
Rohl et al., Protein Structure Prediction Using Rosetta, Methods in Enzymology, vol. 383, Jan. 1, 2004, pp. 66-93. [cited by applicant]
Shoemake, Animating Rotation with Quaternion Curves, Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques, vol. 19, No. 3, Jul. 1, 1985, pp. 245-254. [cited by applicant]
Sohl-Dickstein et al., Deep Unsupervised Learning using Nonequilibrium Thermodynamics, International Conference on Machine Learning, Jun. 1, 2015, pp. 2256-2265, arXiv:1503.03585v8 [cs.LG]. [cited by applicant]
Strokach et al., Deep generative modeling for protein design, ScienceDirect, Current Opinion in Structural Biology, 72, Feb. 1, 2022, pp. 226-236. [cited by applicant]
Strokach et al., Fast and Flexible Protein Design Using Deep Graph Neural Networks, Cell Systems, 11, Oct. 21, 2020, pp. 402-411; https://doi.org/10.1016/j.cels.2020.08.016. [cited by applicant]
Xu et al., GeoDiff: A Geometric Diffusion Model for Molecular Conformation Generation, International Conference on Learning Representations, Mar. 6, 2022, pp. 1-19, arXiv:2023.02923v1 [cs.LG]. [cited by applicant]