IP Library › Granted Patent US 12,694,946
Granted Patent B2
US 12,694,946 · App. 18/240,334 · Granted Jul 28, 2026

Machine learning model distillation for protein design

Inventors: Igor Melnyk (White Plains, NY); Aurelie Chloe Lozano (Scarsdale, NY); Payel Das (New York, NY); Enara C. Vijil (Millwood, NY)
Assignee: International Business Machines Corporation
G16B15/20G06N20/00G16B15/30G16B40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,946
App. No.
18/240,334
Granted
Jul 28, 2026
Kind
B2
Abstract

A distilled machine learning model is produced via initializing a first model with initial weights. An input protein sequence is input into both the first model and a folding protein model, wherein the inputting to the first model generates logits, and wherein the inputting to the folding protein model generates one or more predictive metrics. The one or more predictive metrics are discretized into classes and a first cross-entropy loss is computed based on the logits and the classes. The first model is optimized based on the first cross-entropy loss so that the optimized first model is the distilled machine learning model and an additional machine learning model is trained, using the distilled machine learning model, to perform a downstream protein modeling task.

Claims (68)

1 . A computer-implemented method comprising:

producing a distilled machine learning model via:

initializing a first model with initial weights;

inputting an input protein sequence into both the first model and a folding protein model, wherein the inputting to the first model generates logits, and wherein the inputting to the folding protein model generates one or more predictive metrics;

discretizing the one or more predictive metrics into classes;

computing a first cross-entropy loss based on the logits and the classes; and

optimizing the first model based on the first cross-entropy loss so that the optimized first model is the distilled machine learning model; and

training, using the distilled machine learning model, an additional machine learning model to perform a downstream protein modeling task.

2 . The method of claim 1 , further comprising:

producing a protein design using the trained additional machine learning model; and

synthesizing a protein based on the protein design.

3 . The method of claim 2 , wherein the training of the additional machine learning model comprises:

inputting, separately, protein structures into the additional machine learning model, wherein, in response, the additional machine learning model generates as output a respective predicted protein sequence corresponding to the input protein structures.

4 . The method of claim 1 , wherein the training of the additional machine learning model comprises:

feeding output from the additional machine learning model to the distilled machine learning model, wherein the distilled machine learning model, in response, generates a structure consistency score; and

optimizing the additional machine learning model based on a second cross entropy loss and on a loss of the structure consistency score, wherein the second cross entropy loss is based on a ground truth value of the output of the additional machine learning model.

5 . The method of claim 4 , wherein the output of the additional machine learning model comprises a protein design.

6 . The method of claim 4 , wherein the downstream protein modeling task comprises an inverse protein folding task and the structure consistency score regularizes the inverse protein folding task.

7 . The method of claim 4 , wherein the downstream protein modeling task comprises a protein infilling task and the structure consistency score regularizes the protein infilling task.

8 . The method of claim 1 , wherein the folding protein model:

infers protein structure based on a protein sequence that is input,

comprises a deep learning model and an attention network,

was trained from a public repository of protein sequences and structures, and

is larger than the distilled machine learning model.

9 . The method of claim 1 , wherein the producing a distilled machine learning model further comprises:

generating a set of one or more metrics via a ground truth three-dimensional structure; and

discretizing the set of the one or more metrics into the classes such that the cross-entropy loss is further based on the one or more metrics.

10 . A computer program product, comprising:

one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising:

producing a distilled machine learning model via:

initializing a first model with initial weights;

inputting an input protein sequence into both the first model and a folding protein model, wherein the inputting to the first model generates logits, and wherein the inputting to the folding protein model generates one or more predictive metrics;

discretizing the one or more predictive metrics into classes;

computing a first cross-entropy loss based on the logits and the classes; and

optimizing the first model based on the first cross-entropy loss so that the optimized first model is the distilled machine learning model; and

training, using the distilled machine learning model, an additional machine learning model to perform a downstream protein modeling task.

11 . The computer program product of claim 10 , the instructions further comprising:

producing a protein design using the trained additional machine learning model; and

synthesizing a protein based on the protein design.

12 . A system comprising:

a memory; and

at least one processor, coupled to said memory, and operative to perform operations comprising:

producing a distilled machine learning model via:

initializing a first model with initial weights;

inputting an input protein sequence into both the first model and a folding protein model, wherein the inputting to the first model generates logits, and wherein the inputting to the folding protein model generates one or more predictive metrics;

discretizing the one or more predictive metrics into classes;

computing a first cross-entropy loss based on the logits and the classes; and

optimizing the first model based on the first cross-entropy loss so that the optimized first model is the distilled machine learning model; and

training, using the distilled machine learning model, an additional machine learning model to perform a downstream protein modeling task.

13 . The system of claim 12 , the operations further comprising:

producing a protein design using the trained additional machine learning model; and

synthesizing a protein based on the protein design.

14 . The system of claim 13 , wherein the training of the additional machine learning model comprises:

inputting, separately, protein structures into the additional machine learning model, wherein, in response, the additional machine learning model generates as output a respective predicted protein sequence corresponding to the input protein structures.

15 . The system of claim 12 , wherein the training of the additional machine learning model comprises:

feeding output from the additional machine learning model to the distilled machine learning model, wherein the distilled machine learning model, in response, generates a structure consistency score; and

optimizing the additional machine learning model based on a second cross entropy loss and on a loss of the structure consistency score, wherein the second cross entropy loss is based on a ground truth value of the output of the additional machine learning model.

16 . The system of claim 15 , wherein the output of the additional machine learning model comprises a protein design.

17 . The system of claim 15 , wherein the downstream protein modeling task comprises an inverse protein folding task and the structure consistency score regularizes the inverse protein folding task.

18 . The system of claim 15 , wherein the downstream protein modeling task comprises a protein infilling task and the structure consistency score regularizes the protein infilling task.

19 . The system of claim 12 , wherein the folding protein model:

infers protein structure based on a protein sequence that is input,

comprises a deep learning model and an attention network,

was trained from a public repository of protein sequences and structures, and

is larger than the distilled machine learning model.

20 . The system of claim 12 , wherein the producing a distilled machine learning model further comprises:

generating a set of one or more metrics via a ground truth three-dimensional structure; and

discretizing the set of the one or more metrics into the classes such that the cross-entropy loss is further based on the one or more metrics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2023
From: MELNYK, IGOR; LOZANO, AURELIE CHLOE; DAS, PAYEL; VIJIL, ENARA C.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064759/0286 →
Continuity (1)
Related Publication 20250078953A1 · Mar 6, 2025
References Cited (39)
US 11545238B2 · Greving · 2023 [cited by examiner]
US 11748568B1 · Orhan · 2023 [cited by examiner]
US 11783227B2 · Li · 2023 [cited by examiner]
US 12002094B2 · Kamkar · 2024 [cited by examiner]
US 12008452B2 · Nushi · 2024 [cited by examiner]
US 12211588B1 · Aithani · 2025 [cited by examiner]
US 12462903B2 · Tonekaboni · 2025 [cited by examiner]
US 20150356461A1 · Vinyals · 2015 [cited by examiner]
US 20210027860A1 · Zhang · 2021 [cited by applicant]
US 20210209499A1 · Mishraky · 2021 [cited by examiner]
US 20210304003A1 · Johnson · 2021 [cited by examiner]
US 20210350819A1 · Wu · 2021 [cited by examiner]
US 20220092366A1 · Chiu · 2022 [cited by examiner]
US 20220366263A1 · Ji · 2022 [cited by examiner]
US 20220375538A1 · Das · 2022 [cited by applicant]
US 20230031691A1 · Carroll · 2023 [cited by examiner]
US 20230245722A1 · Liszka · 2023 [cited by examiner]
US 20230325676A1 · Zhang · 2023 [cited by examiner]
US 20250006306A1 · Hoang · 2025 [cited by examiner]
US 20250078953A1 · Melnyk · 2025 [cited by examiner]
US 20250118391A1 · Yuan · 2025 [cited by examiner]
US 20250322902A1 · Ramanathan · 2025 [cited by examiner]
US 20250378364A1 · Tayeb · 2025 [cited by examiner]
WO 2022167325A1 · 2022 [cited by applicant]
C. Hsu et a.l, “Learning inverse folding from millions of predicted structures,” bioRxiv, Posted Sep. 6, 2022. https://doi.org/10.1101/2022.04.10.487779 22 pages. [cited by applicant]
I. Melnyk et al., “AlphaFold Distillation for Improved Inverse Protein Folding,” [Submitted on Oct. 5, 2022], https://arxiv.org/abs/2210.03488 Grace Period Disclosure 20 pages. [cited by applicant]
K. Yang et al., “Masked Inverse Folding with Sequence Transfer for Protein Representation Learning,” bioRxiv, https://doi.org/10.1101/2022.05.25.493516, Posted Mar. 19, 2023. 18 pages. [cited by applicant]
M. Jendrusch et al., AlphaDesign: A de novo protein design framework based on AlphaFold. bioRxiv, Posted Oct. 12, 2021, https://doi.org/10.1101/2021.10.11.463937 20 pages. [cited by applicant]
P. Das, The Promise of Foundation Models and Generative AI, http://helper.ipam.ucla.edu/publications/ems2023/ems2023_18216.pdf 40 pages. [cited by applicant]
Z. Gao et al., “PiFold: Toward effective and efficient protein inverse folding,” [Submitted on Sep. 22, 2022 (v1), last revised Jan. 12, 2023 (this version, v3)], https://arxiv.org/abs/2209.12643 14 pages. [cited by applicant]
Deep Mind, AlphaFold, 16 pages downloaded Jun. 27, 2023 from https://www.deepmind.com/research/highlighted-research/alphafold. [cited by applicant]
C. Hsu et a.l, “Learning inverse folding from millions of predicted structures,” bioRxiv, Posted Sep. 6, 2022. https://doi.org/10.1101/2022.04.10.487779. [cited by applicant]
I. Melnyk et al., “AlphaFold Distillation for Improved Inverse Protein Folding,” [Submitted on Oct. 5, 2022], https://arxiv.org/abs/2210.03488. [cited by applicant]
K. Yang et al., “Masked Inverse Folding with Sequence Transfer for Protein Representation Learning,” bioRxiv, https://doi.org/10.1101/2022.05.25.493516, Posted Mar. 19, 2023. [cited by applicant]
M. Jendrusch et al., AlphaDesign: A de novo protein design framework based on AlphaFold. bioRxiv, Posted Oct. 12, 2021, https://doi.org/10.1101/2021.10.11.463937. [cited by applicant]
P. Das, The Promise of Foundation Models and Generative AI, http://helper.ipam.ucla.edu/publications/ems2023/ems2023_18216.pdf. [cited by applicant]
Z. Gao et al., “PiFold: Toward effective and efficient protein inverse folding,” [Submitted on Sep. 22, 2022 (v1), last revised Jan. 12, 2023 (this version, v3)], https://arxiv.org/abs/2209.12643. [cited by applicant]
Alpha Fold can accurately predict 3D models of protein structures and is accelerating research in nearly every field of biology dowloaded from: https://www.deepmind.com/research/highlighted-research/alphafold Jun. 27, 2… [cited by applicant]
Peter Mell and Timothy Grance, The NIST Definition of Cloud Computing, NIST Special Publication 800-145, Sep. 2011, cover, pp. i-iii and 1-3. [cited by applicant]