IP Library Granted Patent US 12,176,071
Granted Patent B2
US 12,176,071 · App. 16/345,367 · Granted Dec 24, 2024

Systems and methods for ultra-fast identification and abundance estimates of microorganisms using a kmer-depth based approach and privacy-preserving protocols

Inventors: Niamh B. O'Hara (Ridgewood, NY); Rachid Ounit (Riverside, CA)
Assignee: The Joan & Irwin Jacobs Technion-Cornell Institute
G16B30/20C12Q1/6888G16B5/00G16B30/00G16B30/10G16B50/00G16B50/10G16B50/30G16B50/40G16B50/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,176,071
App. No.
16/345,367
Granted
Dec 24, 2024
Kind
B2
Abstract

The present disclosure relates to the use of next generation technologies and analysis using a k-mer based approach which is depth-informed to classify microorganism and estimate abundance in single or mixed microorganismal populations in a rapid manner. The disclosure relates to methods for identifying taxa from environmentally or patient collected samples and processing these samples in a privacy-preserving manner.

Claims (55)

1. A method, comprising:

i) a process of building a database by a data owner, comprising

building a reference k-mer database of taxa specific markers and taxa unique markers; and

setting up the reference k-mer database in at least one server wherein each server stores at least a portion of the database;

ii) a process of searching the database by a user, comprising

obtaining a sample including one or more microorganism populations;

wherein the sample is obtained from a subject;

generating, by a control unit including a memory and a processor, nucleic acid sequence data for the one or more microorganism populations;

determining, by the control unit, a set of k-mers of one or more nucleic acid regions from the one or more microorganism populations;

communicating to the at least one server and comparing, by the control unit, the set of k-mers to the reference k-mer database of taxa-specific markers and taxa unique markers, wherein the at least one server comprises a private or trusted server or a public server or cloud-based server, wherein, when the at least one sever comprises a public server or cloud-based server, the communicating step is under a privacy-preserving scheme;

assigning each k-mer to a taxon only if it is impossible to assign the k-mer to another taxon with the same mapping quality or assignment quality;

filtering, by the control unit, the nucleic acid sequence data that do not map unambiguously to one and only one organism to provide a set of taxa-specific and taxa-unique k-mers found in unambiguously mapped sequence reads, wherein the filtering does not cause loss of identification performance or accuracy;

determining, by the control unit, depth of sequence coverage of taxa specific sequences to identify one or more taxa from the one or more microorganism populations;

modeling, by the control unit, the frequency of k-mers in the sequenced sample that match the database using one or several probabilistic models per taxon; and

modeling, by the control unit, the distribution of the frequency of the taxa-specific and taxa-unique k-mers found in unambiguously mapped sequence reads using one or more probabilistic models to estimate abundance of each identified one or more taxa, wherein the model utilizes a Poisson distribution or a mixture of probabilistic distribution from counts of occurrences of the taxa-specific and taxa-unique k-mers found in the unambiguously mapped reads to estimate the abundance of the one or more taxa; thereby identifying the one or more microorganism populations in the samples; and

iii) administering to the subject a suitable treatment comprising an empiric antimicrobial therapy including one or more of antibiotics, antivirals, antifungals, antiprotozoals, and antihelminthics that appropriately eliminates the one or more microorganism populations identified in the sample.

2. The method of claim 1 , wherein the probabilistic model increases confidence.

3. The method of claim 1 , wherein the set of k-mers includes at least 10 individual k-mers.

4. The method of claim 1 , wherein the set of k-mers includes at least 100 individual k-mers.

5. The method of claim 1 , wherein a subset of the reference database is used that corresponds to the sample source, whether patient or environmentally collected, to increase confidence while maximizing analysis speed.

6. The method of claim 1 , wherein the set of k-mers includes a plurality of k-mers each ranging from about 17 to about 63 nucleotides in length.

7. The method of claim 1 , wherein the set of k-mers includes a plurality of k-mers each ranging from about 17 to about 31 nucleotides in length.

8. The method of claim 1 , wherein the obtained sample comprises DNA or RNA.

9. The method of claim 1 , wherein the nucleic acid sequence(s) from which the nucleic acid sequence data are obtained may be amplified prior to determining a k-mer set to use in analysis.

10. The method of claim 1 , wherein the privacy-preserving scheme comprises at least one of a private information retrieval protocol, an obfuscation protocol, a private set intersection protocol, a homomorphic encryptions scheme, and Shamir secret sharing.

11. The method of claim 1 , wherein the process in step ii) further comprises

comparing the identified one or more taxa in the sample to taxa in other geographic and temporal samples and/or in the context of potential outbreaks of pathogens and antimicrobial resistance, thereby interpreting the identified one or more taxa in a broader public health context.

12. The method of claim 1 , wherein the process in step ii) further comprises

transmitting the identified one or more taxa in the sample and/or actionable recommendations to at least one relevant entity comprising a First Responder, physicians, public health personnel, and/or law enforcement.

13. The method of claim 1 , wherein the one or more microorganism populations in the sample comprises one or more of Clostridium difficile, Staphylococcus aureus (including Methicillin resistant strains (MRSA)), Klebsiella pneumoniae, Klebsiella oxytoca, Escherichia coli, Enterococcus species (including E. faecalis and E. faecium and Vancomycin-resistant Enteroccus (VRE)), Pseudomonas aeruginosa, Candida species (including C. albicans, C. parapsilosis , and C. glabrata ), Streptococcus species, Coagulase-negative staphylococcus species, Enterobacter species, Acinetobacter baumannii, Proteus mirabilis, Stenotrophomonas maltophilia, Citrobacter species, Serratia species, Bacteroides species, Haemophilus species, adenovirus, herpes simplex virus, parainfluenza virus and norovirus, Peptostreptococcus species, additional Klebsiella species, additional Clostridium species, Prevotella species, Morganella morganii, Lactobacillus species, Avian influenza virus, swine influenza virus, Zika virus, West Nile virus, Ebola virus, plasmodium species, yellow fever virus, Dengue virus, Lassa virus, Human immunodeficiency virus, Vibrio cholerae , Middle East respiratory syndrome coronavirus (MERS-COV), Mycobacterium tuberculosis, Borrelia burgdorferi, Yersinia pestis , or a virus causative of viral hepatitis.

14. A method, comprising:

i) a process of building a database by a data owner, comprising

building a reference k-mer database of taxa specific markers and taxa unique markers; and

setting up the reference k-mer database in at least one server, wherein each server stores at least a portion of the database; and

ii) a process of searching the database by a user, comprising

obtaining a sample at a location including one or more microorganism populations at one or more time points

wherein the sample is obtained from a subject, or the sample is an environmental sample;

generating, by a control unit including a memory and a processor, nucleic acid sequence data for the one or more microorganism populations at each of the one or more time points;

determining, by the control unit, a set of k-mers of one or more nucleic acid regions from the one or more microorganism populations;

communicating to the at least one server and comparing, by the control unit, the set of k-mers to the reference k-mer database of taxa-specific markers and taxa unique markers, wherein the at least one server comprises a private or trusted server or a public server or cloud-based server, wherein, when the at least one server comprises a public server or cloud-based server, the communicating step is under a privacy-preserving scheme;

assigning each k-mer to a taxon only if it is impossible to assign the k-mer to another taxon with the same mapping quality or assignment quality;

filtering, by the control unit, the nucleic acid sequence data that do not map unambiguously to one and only one organism to provide a set of taxa-specific and taxa-unique k-mers found in unambiguously mapped sequence reads, wherein the filtering does not cause loss of identification performance or accuracy;

determining, by the control unit, depth of sequence coverage of taxa specific sequences to identify one or more taxa from the one or more microorganism populations;

modeling, by the control unit, the frequency of k-mers in the sequences sample that match the database using one or several probabilistic models per taxon;

modeling, by the control unit, the distribution of frequency of the taxa-specific and taxa-unique k-mers found in unambiguously mapped sequence reads using one or more probabilistic models to estimate abundance of each identified one or more taxa, wherein the model utilizes a Poisson distribution or a mixture of probabilistic distribution from counts of occurrences of the taxa-specific and taxa-unique k-mers found in the unambiguously mapped reads to estimate the abundance of the one or more taxa;

characterizing, based on the identified taxa, the pathogenicity of the one or more microorganism populations; and

iii) administering to the subject a suitable treatment comprising an empiric antimicrobial therapy including one or more of antibiotics, antivirals, antifungals, antiprotozoals, and antihelminthics for eliminating the one or more microorganism populations based on the pathogenicity characterization; or cleaning the environmental source of the sample with an appropriately adjusted cleaning approach comprising specific cleaning products that appropriately eliminates the one or more microorganism populations based on the pathogenicity characterization rather than cleaning blindly.

15. The method of claim 14 , wherein the probabilistic model increases confidence.

16. The method of claim 14 , wherein the set of k-mers includes at least 10 individual k-mers.

17. The method of claim 14 , wherein the set of k-mers includes at least 100 individual k-mers.

18. The method of claim 14 , wherein the set of k-mers includes a plurality of k-mers each ranging from about 17 to about 63 nucleotides in length.

19. The method of claim 14 , wherein the set of k-mers includes a plurality of k-mers each ranging from about 17 to about 31 nucleotides in length.

20. The method of claim 14 , wherein a subset of the reference database is used that corresponds to the sample source, whether patient or environmentally collected, to increase confidence while maximizing analysis speed.

21. The method of claim 14 , wherein the obtained sample comprises DNA or RNA.

22. The method of claim 14 , wherein the one or more microorganism populations in the sample comprises one or more of Clostridium difficile, Staphylococcus aureus (including Methicillin resistant strains (MRSA)), Klebsiella pneumoniae, Klebsiella oxytoca, Escherichia coli, Enterococcus species (including E. faecalis and E. faecium and Vancomycin-resistant Enteroccus (VRE)), Pseudomonas aeruginosa, Candida species (including C. albicans, C. parapsilosis , and C. glabrata ), Streptococcus species, Coagulase-negative staphylococcus species, Enterobacter species, Acinetobacter baumannii, Proteus mirabilis, Stenotrophomonas maltophilia, Citrobacter species, Serratia species, Bacteroides species, Haemophilus species, adenovirus, herpes simplex virus, parainfluenza virus and norovirus, Peptostreptococcus species, additional Klebsiella species, additional Clostridium species, Prevotella species, Morganella morganii, Lactobacillus species, Avian influenza virus, swine influenza virus, Zika virus, West Nile virus, Ebola virus, plasmodium species, yellow fever virus, Dengue virus, Lassa virus, Human immunodeficiency virus, Vibrio cholerae , Middle East respiratory syndrome coronavirus (MERS-COV), Mycobacterium tuberculosis, Borrelia burgdorferi, Yersinia pestis , or a virus causative of viral hepatitis.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2024
From: O'HARA, NIAMH B.; OUNIT, RACHID
To: THE JOAN & IRWIN JACOBS TECHNION-CORNELL INSTITUTE
Reel/Frame 068946/0642 →
Continuity (1)
Related Publication 20190318807A1 · Oct 17, 2019