IP Library › Granted Patent US 12,725,679
Granted Patent B2
US 12,725,679 · App. 18/345,318 · Granted Sep 1, 2026

Generative modeling and representational learning from multi-sequence alignment and phylogenetic tree data

Inventors: Thanh Lam Hoang (Maynooth, IE); Marcos Martínez Galindo (Dublin, IE); Gabriele Picco (Dublin, IE); Mykhaylo Zayats (Dublin, IE); Nhan Huu Pham (Tarrytown, NY); Lam Minh Nguyen (Ossining, NY); Marco Luca Sbodio (Dublin, IE); Dzung Tien Phan (Pleasantville, NY); Vanessa Lopez Garcia (Dublin, IE)
Assignee: International Business Machines Corporation
G16B40/00G16B10/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,679
App. No.
18/345,318
Granted
Sep 1, 2026
Kind
B2
Abstract

Generative modeling from phylogenetic data is provided. The method comprises creating a multi-sequence alignment (MSA) based on a nucleic acid or protein sequence and generating a phylogenetic tree based on the MSA. The phylogenetic tree is fed into a number of machine learning models, which generate vector representations of the nucleic acid or protein sequences based on the phylogenetic tree. The machine learning models generate from the vector representation predicted nucleic acid or protein sequences for at least one of an evolution sequence, regression sequence, or sibling sequences of nucleic acids or proteins according to the phylogenetic tree.

Claims (61)

1 . A computer-implement method of generative modeling from phylogenetic data, the method comprising:

using a number of processors to perform:

creating a multi-sequence alignment (MSA) based on a nucleic acid or protein sequence;

generating a phylogenetic tree based on the MSA;

feeding the phylogenetic tree into a number of machine learning models;

generating, by the machine learning models, vector representations of the nucleic acid or protein sequences based on the phylogenetic tree; and

generating, by the machine learning models from the vector representation, complete predicted nucleic acid or protein sequences for at least one of an evolution sequence, regression sequence, or sibling sequences of nucleic acids or proteins according to the phylogenetic tree.

2 . The method of claim 1 , wherein the machine learning models comprise at least one of:

transformers with condition generation heads; or

transformers with masked language model heads.

3 . The method of claim 1 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and minimize cross entropy loss of sequence generation tasks.

4 . The method of claim 1 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and ignore tokens in the output sequence that have matching tokens in the input nucleic acid or protein sequence.

5 . The method of claim 1 , further comprising creating training data for predicting sibling sequences by:

finding siblings of a given leaf node of the phylogenetic tree; and

organizing the siblings and the given leaf node into a number of different pairs having alternate sequences.

6 . The method of claim 1 , further comprising creating training data for predicting regression sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is first in sequence in the pair.

7 . The method of claim 1 , further comprising creating training data for predicting evolution sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is second in sequence in the pair.

8 . A system for generative modeling from phylogenetic data, the system comprising:

a storage device that stores program instructions;

one or more processors operably connected to the storage device and configured to execute the program instructions to cause the system to:

create a multi-sequence alignment (MSA) based on a nucleic acid or protein sequence;

generate a phylogenetic tree based on the MSA;

feed the phylogenetic tree into a number of machine learning models;

generate, by the machine learning models, vector representations of the nucleic acid or protein sequences based on the phylogenetic tree; and

generate, by the machine learning models from the vector representation, complete predicted nucleic acid or protein sequences for at least one of an evolution sequence, regression sequence, or sibling sequences of nucleic acids or proteins according to the phylogenetic tree.

9 . The system of claim 8 , wherein the machine learning models comprise at least one of:

transformers with condition generation heads; or

transformers with masked language model heads.

10 . The system of claim 8 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and minimize cross entropy loss of sequence generation tasks.

11 . The system of claim 8 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and ignore tokens in the output sequence that have matching tokens in the input nucleic acid or protein sequence.

12 . The system of claim 8 , wherein the program instructions further cause the system to create training data for predicting sibling sequences by:

finding siblings of a given leaf node of the phylogenetic tree; and

organizing the siblings and the given leaf node into a number of different pairs having alternate sequences.

13 . The system of claim 8 , wherein the program instructions further cause the system to create training data for predicting regression sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is first in sequence in the pair.

14 . The system of claim 8 , wherein the program instructions further cause the system to create training data for predicting evolution sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is second in sequence in the pair.

15 . A computer program product for generative modeling from phylogenetic data, the computer program product comprising:

a persistent storage medium having program instructions configured to cause one or more processors to:

create a multi-sequence alignment (MSA) based on a nucleic acid or protein sequence;

generate a phylogenetic tree based on the MSA;

feed the phylogenetic tree into a number of machine learning models;

generate, by the machine learning models, vector representations of the nucleic acid or protein sequences based on the phylogenetic tree; and

generate, by the machine learning models from the vector representation, complete predicted nucleic acid or protein sequences for at least one of an evolution sequence, regression sequence, or sibling sequences of nucleic acids or proteins according to the phylogenetic tree.

16 . The computer program product of claim 15 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and minimize cross entropy loss of sequence generation tasks.

17 . The computer program product of claim 15 , wherein, during training, the machine learning models are provided an input nucleic acid or protein sequence and an output nucleic acid or protein sequence and ignore tokens in the output sequence that have matching tokens in the input nucleic acid or protein sequence.

18 . The computer program product of claim 15 , wherein the program instructions are further configured to cause the processors to create training data for predicting sibling sequences by:

finding siblings of a given leaf node of the phylogenetic tree; and

organizing the siblings and the given leaf node into a number of different pairs having alternate sequences.

19 . The computer program product of claim 15 , wherein the program instructions are further configured to cause the processors to create training data for predicting regression sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is first in sequence in the pair.

20 . The computer program product of claim 15 , wherein the program instructions are further configured to cause the processors to create training data for predicting evolution sequences by:

for a given leaf node in the phylogenetic tree, finding the closest leaf node that is a sibling of a parent node of the given node in the phylogenetic tree; and

pairing the given leaf node with the closest leaf node that is a sibling of the parent node, wherein the given leaf node is second in sequence in the pair.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: HOANG, THANH LAM; MARTÍNEZ GALINDO, MARCOS; PICCO, GABRIELE; ZAYATS, MYKHAYLO; PHAM, NHAN HUU; NGUYEN, LAM MINH; SBODIO, MARCO LUCA; PHAN, DZUNG TIEN; LOPEZ GARCIA, VANESSA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064128/0446 →
Continuity (1)
Related Publication 20250006306A1 · Jan 2, 2025
References Cited (31)
US 20130217006A1 · Sorenson et al. · 2013 [cited by applicant]
US 20210174246A1 · Triplet · 2021 [cited by applicant]
US 20210174903A1 · Rothberg et al. · 2021 [cited by applicant]
US 20210174906A1 · Ul Ain · 2021 [cited by examiner]
US 20210287088A1 · Peng et al. · 2021 [cited by applicant]
US 20220392566A1 · Chatterjee et al. · 2022 [cited by applicant]
US 20230207054A1 · Rashid · 2023 [cited by examiner]
US 20230307092A1 · Senapathy · 2023 [cited by examiner]
CN 112926629A · 2021 [cited by applicant]
CN 112949850A · 2021 [cited by applicant]
CN 113269322A · 2021 [cited by applicant]
CN 113723615A · 2021 [cited by applicant]
CN 114114911A · 2022 [cited by applicant]
KR 1020220151257A · 2022 [cited by applicant]
WO 2022022816A1 · 2022 [cited by applicant]
WO 2022112255A1 · 2022 [cited by applicant]
Cardona (Comparison of Tree-Child Phylogenetic Networks, IEEE, 2009, pp. 552-569) (Year: 2009). [cited by examiner]
Ding et al., “Deciphering protein evolution and fitness landscapes with latent space models”, Nature Communications, vol. 10, Article 5644, 2019, pp. 1-13, accessed Jun. 28, 2023, https://www.nature.com/articles/s41467-… [cited by applicant]
Gascuel et al., “Predicting the Ancestral Character Changes in a Tree is Typically Easier than Predicting the Root State”, Oxford Academic, Systematic Biology, vol. 63, Issue 3, May 2014, pp. 421-435, accessed Jun. 28, … [cited by applicant]
Layer et al., “Phylogenetic trees and Euclidean embeddings”, Arxiv.org, May 3, 2016, pp. 1-12, accessed Jun. 28, 2023, https://arxiv.org/abs/1605.01039. [cited by applicant]
Matsumoto et al, “Novel metric for hyperbolic phylogenetic tree embeddings”, Oxford Academic, Biology Methods & Protocols, vol. 6, Issue 1, 2021, pp. 1-11, accessed Jun. 28, 2023, https:// academic.oup.com/biomethods/ar… [cited by applicant]
Pagel, “The Maximum Likelihood Approach to Reconstructing Ancestral Character States of Discrete Characters on Phylogenies”, Oxford Academic, Systematic Biology, vol. 48, Issue 3, Jul. 1, 1999, pp. 612-622, accessed Jun… [cited by applicant]
Rao et al., “MSA Transformer”, bioRxiv, Feb. 13, 2021, pp. 1-16, accessed Jun. 28, 2023, https://www.biorxiv.org/content/10.1101/2021.02.12.430858v1. [cited by applicant]
Salama et al, “The prediction of virus mutation using neural networks and rough set techniques”, EURASIP Journal on Bioinformatics and Systems Biology, May 13, 2016, pp. 1-11, accessed Jun. 28, 2023, https://bsb-eurasip… [cited by applicant]
Steipe et al., “Sequence Statistics Reliably Predict Stabilizing Mutations in a Protein Domain”, Science Direct, Journal of Molecular Biology, vol. 240, Issue 3, Jul. 14, 1994, pp. 188-192, accessed Jun. 28, 2023, https… [cited by applicant]
Zhu et al., “Applying Neural Network to Reconstruction of Phylogenetic Tree”, ICMLC 2021: 2021 13th International Conference on Machine Learning and Computing, Feb. 2021, pp. 146-152, accessed Jun. 28, 2023, https://di.… [cited by applicant]
Github. “Generative Toolkit 4 Scientific Discovery”, retrieved from web https://web.archive.org/web/20230514071708/https://github.com/GT4SD, May 14, 2023, 4 pages. [cited by applicant]
Wikipedia. “Diagram”, retrieved from web https://web.archive.org/web/20230507084252/https://en.wikipedia.org/wiki/Diagram, May 7, 2023, 7 pages. [cited by applicant]
Wikipedia. “Evolution”, retrieved from web https://web.archive.org/web/20230504012905/https://en.wikipedia.org/wiki/Evolution, May 4, 2023, 83 pages. [cited by applicant]
Wikipedia. “Species”, retrieved from web https://web.archive.org/web/20230511014806/https://en.wikipedia.org/wiki/Species, May 11, 2023, 31 pages. [cited by applicant]
Wikipedia. “Tree (graph theory)”, retrieved from web https://web.archive.org/web/20230503032335/https://en.wikipedia.org/wiki/Tree_(graph_theory), May 3, 2023, 8 pages. [cited by applicant]