IP Library Patent Application 17411889
Patent Application
App. No. 17/411,889

SYSTEMS AND METHODS FOR NUCLEIC ACID SEQUENCE ASSEMBLY

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/411,889
Abstract

Methods, processes, and particularly computer implemented processes and computer program products are provided for use in the analysis of genetic sequence data. The processes and products are employed in the assembly of shorter nucleic acid sequence data into longer linked and preferably contiguous genetic constructs, including large contigs, chromosomes and whole genomes.

Claims (60)

1 . A sequencing method of assembling nucleic acid sequence reads comprising:

at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:

obtaining a plurality of sequence reads derived from a larger contiguous nucleic acid, wherein two or more sequence reads derived from a common fragment of the larger contiguous nucleic acid comprise a common barcode sequence,

identifying a first subset of sequence reads in the plurality of sequence reads that comprise both overlapping sequences and a common barcode sequence; and

aligning the first subset of sequence reads to provide a contiguous linear nucleic acid sequence.

2 . The method of claim 1 , further comprising repeating the identifying and aligning steps with a plurality of different subsets of sequence reads to provide a plurality of contiguous linear nucleic acid sequences.

3 . The method of claim 2 , further comprising ordering the plurality of different contiguous linear nucleic acid sequences in a sequence context within the larger contiguous nucleic acid.

4 . The method of claim 3 , wherein the ordering comprises mapping the plurality of different contiguous linear nucleic acid sequences against a reference sequence.

5 . The method of claim 3 , wherein the ordering comprises:

identifying one or more sequence reads that comprise a barcode sequence common to a first contiguous linear nucleic acid sequence, but include overlapping sequences with a second contiguous linear nucleic acid sequence; and

identifying the first and second contiguous linear nucleic acid as structurally linked.

6 . A method of assembling nucleic acid sequence reads into larger contiguous sequences, comprising:

at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:

obtaining a plurality of sequence reads derived from a larger contiguous nucleic acid, identifying a first subsequence from a set of overlapping sequence reads in the plurality of sequence reads;

extending the first subsequence to one or more adjacent or overlapping sequences based upon the presence of a barcode sequence on the adjacent sequence that is common to the first subsequence; and

providing a linear nucleic acid sequence that comprises the first subsequence and the one or more adjacent sequences.

7 . A sequencing method comprising, at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:

(A) obtaining a plurality of sequence reads, wherein

the plurality of sequence reads comprises a plurality of sets of sequence reads,

each respective sequence read in a set of sequence reads includes (i) a first portion that corresponds to a subset of a larger contiguous nucleic acid and (ii) a common second portion that forms an identifier that is independent of the sequence of the larger contiguous nucleic acid and that identifies a partition, in a plurality of partitions, in which the respective sequence read was formed, and

each respective set of sequence reads in the plurality of sets of sequence reads is formed in a partition in the plurality of partitions and each partition includes one or more fragments of the larger contiguous nucleic acid that is used as the template for each respective sequence read in the partition;

(B) creating a respective set of k-mers for each sequence read in the plurality of sequence reads, wherein

the sets of k-mers collectively comprise a plurality of k-mers,

the identifiers of the sequence reads for each k-mer in the plurality of k-mers is retained,

k is less than the average length of the sequence reads in the plurality of sequence reads, and

each respective set of k-mers includes at least eighty percent of the possible k-mers of length k of the first portion of the corresponding sequence read;

(C) tracking, for each respective k-mer in the plurality of k-mers, an identity of each sequence read in the plurality of sequence reads that contains the respective k-mer and the identifier of the set of sequence reads that contains the sequence read;

(D) graphing the plurality k-mers as a graph comprising a plurality of nodes connected by a plurality of directed arcs, wherein

each node comprises an uninterrupted set of k-mers in the plurality of k-mers of length k with k−1 overlap,

each arc connects an origin node to a destination node in the plurality of nodes,

a final k-mer of an origin node has k−1 overlap with an initial k-mer of a destination node, and

a first origin node has a first directed arc with both a first destination node and a second destination node in the plurality of nodes; and

(E) determining whether to merge the origin node with the first destination node or the second destination node in order to derive a contig sequence that is more likely to be representative of a portion of the larger contiguous nucleic acid, wherein the contig sequence comprises (i) the origin node and (ii) one of the first destination node and the second destination node, wherein the determining uses at least the identifiers of the sequence reads for k-mers in the first origin node, the first destination node, and the second destination node.

8 . The sequencing method of claim 7 , wherein,

the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node,

the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and

the determining (E) comprises determining whether the first path is more likely representative of the larger contiguous nucleic acid than the second path by evaluating a number of identifiers shared between the k-mers of the nodes of a first portion of the first path and the k-mers of the nodes of a second portion of the first path versus a number of identifiers shared between the k-mers of the nodes of a first portion of the second path and the k-mers of the nodes of a second portion of the second path.

9 . The sequencing method of claim 7 , wherein

the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node,

the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and

the determining (E) up-weights the first path relative to the second path when the first path has higher average coverage than the second path.

10 . The sequencing method of claim 7 , wherein the

the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node,

the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and

the determining (E) up-weights the first path relative to the second path when the first path represents a longer contiguous portion of the larger contiguous nucleic acid sequence than the second path.

11 . The sequencing method of claim 7 , wherein a first k-mer in the first node is present in a sub-plurality of the plurality of sequence reads and the identity of each sequence read in the sub-plurality of sequence reads is retained for the first k-mer and used by the determining (E) to determine whether the first path is more likely representative of the larger contiguous nucleic acid sequence than the second path.

12 . The sequencing method of claim 7 , wherein

a partition in the plurality of partitions comprises at least 1000 molecules with the common second portion, and

each molecule in the at least 1000 molecules further comprises a primer sequence complementary to at least a portion of the larger contiguous nucleic acid.

13 . The sequencing method of claim 7 , wherein

a partition in the plurality of partitions comprises at least 1000 molecules with the common second portion, and

each molecule in the at least 1000 molecules further comprises a primer site and a semi-random N-mer priming sequence that is complementary to part of the larger contiguous nucleic acid.

14 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions are greater than 50 kilobases in length.

15 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions are between 20 kilobases and 200 kilobases in length.

16 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions consists of between 1 and 500 different fragments of the larger contiguous nucleic acid.

17 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions consists of between 5 and 100 fragments of the larger contiguous nucleic acid.

18 . The sequencing method of claim 7 , wherein the plurality of sequence reads is obtained from less than 5 nanograms of nucleic or ribonucleic acid.

19 . The sequencing method of claim 7 , wherein the identifier in the second portion of each respective sequence read in the set of sequence reads encodes a common value selected from the set {1, . . . , 1024}, the set {1, . . . , 4096}, the set {1, . . . , 16384}, the set {1, . . . , 65536}, the set {1, . . . , 262144}, the set {1, . . . , 1048576}, the set {1, . . . , 4194304}, the set {1, . . . , 16777216}, the set {1, . . . , 67108864}, or the set {1, . . . , 1×10 12 }.

20 . The sequencing method of claim 7 , wherein the identifier is an N-mer, and N is an integer selected from the set {4, . . . , 20}.

21 .- 44 . (canceled)

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2022
From: SCHNALL-LEVIN, MICHAEL; MACCALLUM, IAIN
To: 10X GENOMICS, INC.
Reel/Frame 059858/0080 →