CHARACTERIZING THE GENOME OF INDIVIDUAL CELLS BY LONG FRAGMENT READ SEQUENCING OF OLIGONUCLEOTIDE TAGGED DNA FRAGMENTS
This disclosure provides technology for ordering sequence information derived from one or more target polynucleotides. In one aspect, one or more tiers or levels of fragmentation and aliquoting are generated, after which sequence information is obtained from fragments in a final level or tier. Each fragment in such final tier is from a particular aliquot, which, in turn, is from a particular aliquot of a prior tier, and so on. For every fragment of an aliquot in the final tier, the aliquots from which it was derived at every prior tier is known, or can be discerned. Thus, identical sequences from overlapping fragments from different aliquots can be distinguished and grouped as being derived from the same or different fragments from prior tiers. When the fragments in the final tier are sequenced, overlapping sequence regions of fragments in different aliquots are used to register the fragments so that non-overlapping regions are ordered. In one aspect, this process is carried out in a hierarchical fashion until the one or more target polynucleotides are characterized, e.g. by their nucleic acid sequences, or by an ordering of sequence segments, or by an ordering of single nucleotide polymorphisms (SNPs), or the like.
1 . A library of DNA fragments prepared from the genomes of a plurality of cells, labeled for assembly of sequence reads to determine genomic sequences of single cells within the plurality;
wherein the library comprises fragments of genomes from the plurality of cells, each labeled with a first oligonucleotide tag and a second oligonucleotide tag;
wherein fragments in the library having the same first oligonucleotide tag originated from the genome of the same cell; and
wherein fragments in the library having the same first oligonucleotide tag and the same second oligonucleotide tag more often contain fragment sequences that originated within 100 kb of each other within the genome of the cell compared with fragments having the same first oligonucleotide tag but a different second oligonucleotide tag;
wherein sequence reads from the labeled fragments can be assembled to obtain complete or partial genome sequences of each of the plurality of cells by a process that comprises ordering and assembling sequence reads, whereby reads that contain the same first and second oligonucleotide tag sequences are grouped together.
2 . The DNA library of claim 1 , prepared by a process that comprises:
labeling genomic fragments from the plurality of cells with a first oligonucleotide tag that identifies the single cell from which the fragment was obtained; and
labeling genomic fragments from each cell with a second oligonucleotide tag that identifies a portion of the genome of the single cell from which the fragment was obtained.
3 . The DNA library of claim 1 , prepared by a process that comprises:
separating single cells from the plurality into a tier of first aliquots so that each aliquot contains a maximum of one cell;
labeling genomic DNA in each of the first aliquots with a first oligonucleotide tag that is unique for each first aliquot;
in each of the first aliquots, making fragments of 100 kb in size from genomic DNA contained in the aliquot;
separating fragments made in each of the first aliquots into a tier of second aliquots;
labeling fragments in the second aliquots with a second oligonucleotide tag that is unique for each second aliquot;
pooling fragments bearing both the first oligonucleotide tag and the second oligonucleotide tag from the second aliquots to form said library.
4 . The DNA library of claim 3 , wherein the process of preparing the library comprises replicating DNA in the first aliquots before labeling with the first oligonucleotide tag.
5 . The DNA library of claim 3 , wherein the process of preparing the library comprises replicating DNA in the second aliquots before labeling with the second oligonucleotide tag.
6 . The DNA library of claim 3 , wherein the process of preparing the library comprises making subfragments of 10 to 30 kb in size from DNA in the second aliquots before labeling with the second oligonucleotide tag.
7 . The DNA library of claim 1 , wherein the process of preparing the library comprises labeling DNA by replication using tagged primers.
8 . The DNA library of claim 1 , wherein the plurality of cells are bacteria.
9 . The DNA library of claim 1 , wherein fragments bearing different first oligonucleotide tags and different second oligonucleotide tags are present in a single solution.
10 . The DNA library of claim 1 , wherein fragments bearing different first oligonucleotide tags and different second oligonucleotide tags are present in separate aliquots.
11 . The DNA library of claim 9 , wherein the single solution is prepared by separately labeling genomic DNA from the cells with different oligonucleotide tags in a plurality of aliquots, then pooling the aliquots to form the single solution.
12 . A method of determining genomic sequences of single cells in a sample of cells, comprising:
providing the DNA library of claim 1 prepared from the sample;
obtaining sequence reads from at least some of the labeled fragments in the DNA library; and
determining complete or partial nucleotide sequences of single cells in the sample by a process that comprises ordering and assembling sequence reads whereby reads that contain the same first and second oligonucleotide tag sequences are grouped together.
13 . The method of claim 1 , wherein the sequence reads are obtained from the labeled fragments by a process that comprises sequencing by synthesis.
14 . A library of DNA fragments prepared from the genomes of a plurality of cells, labeled for assembly of sequence reads to determine genomic sequences of single cells within the plurality;
wherein the library comprises fragments of genomes from the plurality of cells, each labeled with an oligonucleotide tag;
wherein fragments in the library having the same oligonucleotide tag more often contain fragment sequences that originated within 100 kb of each other within the genome of one of the plurality of cells compared with fragments having different oligonucleotide tag sequences;
wherein sequence reads from the labeled fragments can be assembled to obtain complete or partial genome sequences of each of the plurality of cells by a process that comprises ordering and assembling sequence reads, whereby reads that contain the same oligonucleotide tag sequences are grouped together.