IP Library Granted Patent US 12,119,089
Granted Patent B2
US 12,119,089 · App. 16/841,238 · Granted Oct 15, 2024

Generation and use of simulated genomic data

Inventors: Agata Foryciarz (Warsaw, PL); Dennis A. Dean, II (Wenham, MA)
Assignee: SEVEN BRIDGES GENOMICS, INC.
G16B45/00G16B5/00G16B20/00G16B20/20G16B50/00G16B50/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,119,089
App. No.
16/841,238
Granted
Oct 15, 2024
Kind
B2
Abstract

Embodiments of the invention utilize a graph-based approach for simulating genomic datasets from large scale populations. Genomic data may be represented as a directed acyclic graph (DAG) that incorporates individual sample data including variant type, position, and zygosity. A simulator may operate on the DAG to generate variant datasets based on probabilistic traversal of the DAG. This probabilistic traversal reflects genomic variant types associated with the subpopulation used to build the DAG, and as a result, the generated variant datasets maintain statistical fidelity to the original sample data.

Claims (41)

1. A method of simulating genomic data using input data representing nucleic-acid sequences obtained by chemical analysis of biological samples obtained from a population of organisms, the population exhibiting a variant frequency, the method comprising the steps of:

a. computationally representing the input data for the population exhibiting the variant frequency as a directed acyclic graph (DAG) data structure comprising a plurality of nodes and edges connecting the nodes, wherein the nodes include an origin node and a terminus node, each node corresponds to a genomic position and a variant type, and the edges have weights corresponding to the occurrences of nodes connected by edges in the input data; and

b. repeatedly:

computationally traversing a sequence of nodes of the DAG data structure starting from the origin node and terminating with the terminus node to create simulated genomic data for the population with maintaining variant frequency of the population for the input data, the traversed nodes corresponding to a first variant type representing a single-nucleotide polymorphism (SNP), a second variation type representing a haplotype and a third variation type representing insertion and deletion of a nucleotide, the traversal being performed probabilistically in accordance with the edge weights; and

(ii) storing, in a database, the simulated genomic data corresponding to the traversal and converting the simulated genomic data into a standard file format; and

(iii) updating the DAG data structure based on determination whether new variants or update counts of variants already present using a hash value assigned to each node by a hash function, the hash value being a unique value determined based on the genomic position and the variant type associated with each node.

2. The method of claim 1 , wherein each node further stores zygosity information.

3. The method of claim 1 , wherein the traversals are performed using a weighted random function.

4. The method of claim 1 , further comprising repeating steps (a) and (b) for a plurality of populations.

5. The method of claim 4 , further comprising the steps of:

computationally representing input data from each of the populations as a separate DAG data structure; and wherein step (b) is performed on the separate DAG data structures to simulate individuals from each population.

6. The method of claim 1 , wherein the computational representation step comprises:

initializing the DAG data structure; and

repeatedly adding entries to the DAG data structures or changing weights stored in the data structure and associated with existing entries by performing union operations on the DAG data structure and a new entry thereto.

7. The method of claim 6 , wherein each node also comprises zygosity information and entries are added according to steps comprising:

reading in a variant pair as a node;

assigning a node hash value to the node, the node hash value being based at least in part on a genomic position and a zygosity of the variant pair associated with the node; and

performing a union operation using the node hash value to determine if the variant pair associated with the node is already in the DAG data structure, and if not, increasing a weight of at least one edge of the DAG data structure that connects to the node.

8. The method of claim 1 wherein the DAG is created using a set of genotypes for each individual organism in the population.

9. The method of claim 1 , wherein the simulated genomic data is a simulated genome, a simulated genome subset, a simulated chromosome, or a list of simulated genomic variants.

10. The method of claim 1 , wherein the DAG data structure does not store variants that are homozogyous for reference sequence.

11. The method of claim 1 , wherein computationally representing the input data as a DAG data structure further comprises calculating haplotypes from the input data, wherein each haplotype is computationally stored in the DAG data structure as a node.

12. The method of claim 11 , wherein the haplotype comprises a set of variants that are statistically correlated with one another.

13. A system for simulating genomic data using input data representing nucleic-acid sequences obtained by chemical analysis of biological samples obtained from a population of organisms, the population exhibiting a variant frequency, the system comprising:

a computer memory;

a graph generator for computationally representing the input data for the population exhibiting the variant frequency as a data structure in the memory encoding a directed acyclic graph (DAG) comprising a plurality of nodes and edges connecting the nodes, wherein the nodes include an origin node and a terminus node, each node specifies a genomic position and a variant type, and the edges have associated weights corresponding to the occurrences of nodes connected by edges in the input data; and

a population simulator for simulating variant datasets by repeatedly:

(i) traversing a sequence of nodes of the DAG data structure starting from the origin node and terminating with the terminus node to create simulated genomic data for the population with maintaining variant frequency of the population for the input data, the traversed nodes corresponding to a first variant type representing a single-nucleotide polymorphism (SNP), a second variation type representing a haplotype and a third variation type representing insertion and deletion of a nucleotide, the traversal being performed probabilistically in accordance with the edge weights; and

(ii) storing, in the computer memory, the simulated genomic data corresponding to the traversal and converting the simulated genomic data into a standard file format,

the graph generator is further configured to update the DAG data structure based on determination whether new variants or update counts of variants already present using a hash value assigned to each node by a hash function, the hash value being a unique value determined based on the genomic position and the variant type associated with each node.

14. The system of claim 13 , wherein each node further comprises zygosity information.

15. The system of claim 13 , wherein the population simulator is configured to traverse the sequence of nodes using a weighted random function.

16. The system of claim 15 , wherein the population simulator is configured to repeat (i) and (ii) for a plurality of populations.

17. The system of claim 16 , wherein the graph generator is configured to computationally represent input data from each of the populations as a separate DAG, and the population simulator is configured to simulate individuals from each separate DAG.

18. The system of claim 13 , wherein the graph generator is configured to initialize the DAG and repeatedly add entries thereto or change weights associated in the data structure with existing entries by performing union operations on the DAG and a new entry.

19. The system of claim 18 , wherein each node also comprises zygosity information and the graph generator is configured to add entries by:

temporarily storing, in the data structure, a variant pair as a node;

assigning a node hash value to the node, the node hash value being based at least in part on a genomic position and a zygosity of the variant pair associated with the node; and

performing a union operation using the node hash value to determine if the variant pair associated with the node is already in the DAG data structure, and if not, increasing a weight of at least one edge of the DAG data structure that connects to the node.

20. The system of claim 13 , wherein the simulated genomic data is a simulated genome, a simulated genome subset, a simulated chromosome, or a list of simulated genomic variants.

21. The system of claim 13 , wherein the graph generator is further configured to not include homozygous reference variants in the DAG data structure.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2025
From: FORYCIARZ, AGATA; DEAN, DENNIS A.
To: SEVEN BRIDGES GENOMICS, INC.
Reel/Frame 072601/0506 →
RELEASE OF SECURITY INTEREST Recorded Aug 2, 2022
From: IMPERIAL FINANCIAL SERVICES B.V.
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 061055/0078 →
RELEASE OF SECURITY INTEREST Recorded May 24, 2022
From: IMPERIAL FINANCIAL SERVICES B.V.
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 060173/0792 →
SECURITY INTEREST Recorded May 24, 2022
From: SEVEN BRIDGES GENOMICS INC.
To: IMPERIAL FINANCIAL SERVICES B.V.
Reel/Frame 060173/0803 →
SECURITY INTEREST Recorded Mar 30, 2022
From: SEVEN BRIDGES GENOMICS INC.
To: IMPERIAL FINANCIAL SERVICES B.V.
Reel/Frame 059554/0165 →
Continuity (2)
Continuation 15383102 · Dec 19, 2016
Related Publication 20200234797A1 · Jul 23, 2020