IP Library Granted Patent US 12,744,107
Granted Patent B2
US 12,744,107 · App. 16/937,578 · Granted Sep 22, 2026

System and method for nucleotide analysis and prediction of disease risk

Inventors: Diane Wu (San Francisco, CA); Poornima Parameswaran (Menlo Park, CA)
Assignee: MIRATERRA INC.
G16B40/00G06N20/00G16B25/00G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,107
App. No.
16/937,578
Granted
Sep 22, 2026
Kind
B2
Abstract

A system and method for the detection of pathogens and other microbes using nucleotide analysis is described. Aligned and unaligned nucleotide sequences are utilized to predict the presence or absence of pathogens and other microbes.

Claims (46)

1 . A computer-implemented method of training a model for prediction of disease risk in soil, comprising:

receiving digitized samples of a plurality of soil samples, the digitized samples including a plurality of sets of nucleic acid sequences of microbes present in the plurality of soil samples, wherein each of the plurality of sets of nucleic acid sequences is associated with a different one of the plurality of soil samples;

determining that at least one of the plurality of sets of nucleic acid sequences includes an unaligned sequence, wherein the unaligned sequence is a nucleic acid sequence that does not align within a threshold number of nucleotides to nucleotides of any known nucleic acid sequences of one or more known microbes predictive of a disease;

determining, for each of the plurality of sets of nucleic acid sequences, whether there is a co-occurrence of (i) a set of nucleic acid sequences of the plurality of sets of nucleic acid sequences including at least the unaligned sequence and (ii) the disease present in a soil sample of the plurality of soil samples associated with the set of nucleic acid sequences;

creating a training data set by associating with the disease the unaligned sequence responsive to determining the co-occurrence for a threshold number of the plurality of sets of nucleic acid sequences; and

training the model using the training data set to predict risk of the disease in a test soil sample, wherein the model accounts for a presence of an ameliorative microbe in the test soil sample that reduces the risk of the disease, and wherein the model learns a predictive effect of co-occurrence of the unaligned sequence and the ameliorative microbe on the risk of the disease.

2 . The method of claim 1 , further comprising:

determining that the unaligned sequence does not correlate to a by-product of the one or more known microbes predictive of the disease.

3 . The method of claim 1 , further comprising:

training the model with metadata describing a location where the plurality of soil samples is obtained.

4 . The method of claim 1 , further comprising:

training the model with metadata including one or more of weather patterns, sources of water, fertilizer use, pesticide use, source of seeds, and operational data about a farm from which the plurality of soil samples is sourced.

5 . The method of claim 1 , further comprising:

determining that the unaligned sequence does not align within a threshold number of nucleotides to the nucleotides of any known nucleic acid sequences by determining absence of a specific loci in the unaligned sequence.

6 . The method of claim 1 , wherein the model is a multi-layered neural network, and wherein the model takes input nucleic acid sequences and outputs phenotypic characteristics.

7 . The method of claim 1 , wherein the disease is citrus greening or strawberry disease.

8 . The method of claim 1 , further comprising:

determining that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to at least one of the nucleotides of one or more known nucleic acid sequences of the one or more known microbes predictive of the disease; and

determining that presence of the different nucleic acid sequence is predictive of the disease.

9 . The method of claim 1 , further comprising:

determining that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to nucleotides of nucleic acid sequences of a microbe known to be a suppressor of at least one disease.

10 . The method of claim 1 , further comprising:

providing an alert regarding a prediction of the model.

11 . The method of claim 1 , wherein the digitized samples are received from a sequencer.

12 . A system for training a model for prediction of disease risk in soil, comprising:

a non-transitory computer-readable storage medium storing instructions, the instructions when executed by one or more processors cause the one or more processors to:

receive digitized samples of a plurality of soil samples, the digitized samples including a plurality of sets of nucleic acid sequences of microbes present in the plurality of soil samples, wherein each of the plurality of sets of nucleic acid sequences is associated with a different one of the plurality of soil samples;

determine that at least one of the plurality of sets of nucleic acid sequences includes an unaligned sequence, wherein the unaligned sequence is a nucleotide sequence that does not align within a threshold number of nucleotides to nucleotides of any known nucleic acid sequences of one or more known microbes predictive of a disease;

determine, for each of the plurality of sets of nucleic acid sequences, whether there is a co-occurrence of (i) a set of nucleic acid sequences of the plurality of sets of nucleic acid sequences including at least the unaligned sequence and (ii) the disease present in a soil sample of the plurality of soil samples associated with the set of nucleic acid sequences;

create a training data set by associating with the disease the unaligned sequence responsive to determining the co-occurrence for a threshold number of the plurality of sets of nucleic acid sequences; and

train the model using the training data set to predict risk of the disease in a test soil sample, wherein the model accounts for a presence of an ameliorative microbe in the test soil sample that reduces the risk of the disease, and wherein the model learns a predictive effect of co-occurrence of the unaligned sequence and the ameliorative microbe on the risk of the disease.

13 . The system of claim 12 , wherein the one or more processors are further configured to:

determine that the unaligned sequence does not correlate to a by-product of the one or more known microbes predictive of the disease.

14 . The system of claim 12 , wherein the one or more processors are further configured to:

train the model with metadata describing a location where the plurality of soil samples is obtained.

15 . The system of claim 12 , wherein the one or more processors are further configured to:

train the model with metadata including one or more of weather patterns, sources of water, fertilizer use, pesticide use, source of seeds, and operational data about a farm from which the plurality of soil samples is sourced.

16 . The system of claim 12 , wherein the one or more processors are further configured to:

determine that the unaligned sequence does not align within a threshold number of nucleotides to the nucleotides of any known nucleic acid sequences by determining absence of a specific loci in the unaligned sequence.

17 . The system of claim 12 , wherein the model is a multi-layered neural network, and wherein the model takes input nucleic acid sequences and outputs phenotypic characteristics.

18 . The system of claim 12 , wherein the disease is citrus greening or strawberry disease.

19 . The system of claim 12 , wherein the one or more processors are further configured to:

determine that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to at least one of the nucleotides of one or more known nucleic acid sequences of the one or more known microbes predictive of the disease; and

determine that presence of the different nucleic acid sequence is predictive of the disease.

20 . The system of claim 12 , wherein the one or more processors are further configured to:

determine that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to nucleotides of nucleic acid sequences of a microbe known to be a suppressor of at least one disease.

Assignments (3)
NUNC PRO TUNC ASSIGNMENT Recorded Jun 18, 2025
From: TRACE GENOMICS, INC.
To: TRACE GENOMICS ABC
Reel/Frame 071453/0501 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2025
From: TRACE GENOMICS ABC
To: MIRATERRA INC.
Reel/Frame 071453/0532 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2020
From: WU, DIANE; PARAMESWARAN, POORNIMA
To: TRACE GENOMICS, INC.
Reel/Frame 053320/0167 →
Continuity (3)
Continuation 15288731 · Oct 7, 2016
Provisional Application 62238615 · Oct 7, 2015
Related Publication 20200357485A1 · Nov 12, 2020
References Cited (28)
US 7005257B1 · Haas et al. · 2006 [cited by applicant]
US 7058616B1 · Larder · 2006 [cited by examiner]
US 10395115B2 · Kumar et al. · 2019 [cited by applicant]
US 20030190603A1 · Larder et al. · 2003 [cited by applicant]
US 20060160071A1 · Heckerman et al. · 2006 [cited by applicant]
US 20070130633A1 · Urban et al. · 2007 [cited by applicant]
US 20120310863A1 · Crockett · 2012 [cited by examiner]
US 20130259899A1 · Allen-Vercoe et al. · 2013 [cited by applicant]
US 20140127718A1 · Ma · 2014 [cited by examiner]
US 20140162274A1 · Kunin et al. · 2014 [cited by applicant]
US 20140213770A1 · Dong et al. · 2014 [cited by applicant]
US 20150056613A1 · Kural · 2015 [cited by examiner]
US 20150267191A1 · Steelman et al. · 2015 [cited by applicant]
US 20150284810A1 · Knight et al. · 2015 [cited by applicant]
US 20160103958A1 · Hebert et al. · 2016 [cited by applicant]
US 20160148104A1 · Itzhaky · 2016 [cited by examiner]
US 20170039316A1 · Fofanov et al. · 2017 [cited by applicant]
US 20170213083A1 · Shriver et al. · 2017 [cited by applicant]
US 20190227046A1 · Parameswaran et al. · 2019 [cited by applicant]
WO WO0070340A2 · 2000 [cited by examiner]
Kennedy et al. (How Flickr Helps US Make Sense of the World: Context and Content in Community-Contributed Media Collections, Sep. 2007, pp. 631-640) (Year: 2007). [cited by examiner]
Nielsen et al. (Microorganisms as indicators of soil health, Jan. 2002, pp. 0-82) (Year: 2002). [cited by examiner]
Cretoiu et al. (A novel salt-tolerant chitobiosidase discovered by genetic screening of a metagenomic library derived from chitin-amended agricultural soil, Apr. 2015, pp. 1-18) (Year: 2015). [cited by examiner]
Van Elsas et al. (The metagenomics of disease-suppressive soils—experiments from the Metacontrol probje4ct, Sep. 2008, pp. 591-601) (Year: 2008). [cited by examiner]
Lawrence, C. E. et al. “An Expectation Maximization (EM) Algorithm for the Identification and Characterization of Common Sites in Unaligned Biopolymer Sequences.” Proteins: Structure, Function, and Genetics, vol. 7, No.… [cited by applicant]
Loewenstern, D. et al. “DNA Sequence Classification Using Compression-Based Induction.” Rutgers University Libraries, May 1995, pp. 1-12. [cited by applicant]
United States Office Action, U.S. Appl. No. 15/288,731, filed Apr. 2, 2019, 22 pages. [cited by applicant]
United States Office Action, U.S. Appl. No. 15/288,731, filed Aug. 30, 2019, 16 pages. [cited by applicant]