System and method for nucleotide analysis and prediction of disease risk
A system and method for the detection of pathogens and other microbes using nucleotide analysis is described. Aligned and unaligned nucleotide sequences are utilized to predict the presence or absence of pathogens and other microbes.
1 . A computer-implemented method of training a model for prediction of disease risk in soil, comprising:
receiving digitized samples of a plurality of soil samples, the digitized samples including a plurality of sets of nucleic acid sequences of microbes present in the plurality of soil samples, wherein each of the plurality of sets of nucleic acid sequences is associated with a different one of the plurality of soil samples;
determining that at least one of the plurality of sets of nucleic acid sequences includes an unaligned sequence, wherein the unaligned sequence is a nucleic acid sequence that does not align within a threshold number of nucleotides to nucleotides of any known nucleic acid sequences of one or more known microbes predictive of a disease;
determining, for each of the plurality of sets of nucleic acid sequences, whether there is a co-occurrence of (i) a set of nucleic acid sequences of the plurality of sets of nucleic acid sequences including at least the unaligned sequence and (ii) the disease present in a soil sample of the plurality of soil samples associated with the set of nucleic acid sequences;
creating a training data set by associating with the disease the unaligned sequence responsive to determining the co-occurrence for a threshold number of the plurality of sets of nucleic acid sequences; and
training the model using the training data set to predict risk of the disease in a test soil sample, wherein the model accounts for a presence of an ameliorative microbe in the test soil sample that reduces the risk of the disease, and wherein the model learns a predictive effect of co-occurrence of the unaligned sequence and the ameliorative microbe on the risk of the disease.
2 . The method of claim 1 , further comprising:
determining that the unaligned sequence does not correlate to a by-product of the one or more known microbes predictive of the disease.
3 . The method of claim 1 , further comprising:
training the model with metadata describing a location where the plurality of soil samples is obtained.
4 . The method of claim 1 , further comprising:
training the model with metadata including one or more of weather patterns, sources of water, fertilizer use, pesticide use, source of seeds, and operational data about a farm from which the plurality of soil samples is sourced.
5 . The method of claim 1 , further comprising:
determining that the unaligned sequence does not align within a threshold number of nucleotides to the nucleotides of any known nucleic acid sequences by determining absence of a specific loci in the unaligned sequence.
6 . The method of claim 1 , wherein the model is a multi-layered neural network, and wherein the model takes input nucleic acid sequences and outputs phenotypic characteristics.
7 . The method of claim 1 , wherein the disease is citrus greening or strawberry disease.
8 . The method of claim 1 , further comprising:
determining that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to at least one of the nucleotides of one or more known nucleic acid sequences of the one or more known microbes predictive of the disease; and
determining that presence of the different nucleic acid sequence is predictive of the disease.
9 . The method of claim 1 , further comprising:
determining that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to nucleotides of nucleic acid sequences of a microbe known to be a suppressor of at least one disease.
10 . The method of claim 1 , further comprising:
providing an alert regarding a prediction of the model.
11 . The method of claim 1 , wherein the digitized samples are received from a sequencer.
12 . A system for training a model for prediction of disease risk in soil, comprising:
a non-transitory computer-readable storage medium storing instructions, the instructions when executed by one or more processors cause the one or more processors to:
receive digitized samples of a plurality of soil samples, the digitized samples including a plurality of sets of nucleic acid sequences of microbes present in the plurality of soil samples, wherein each of the plurality of sets of nucleic acid sequences is associated with a different one of the plurality of soil samples;
determine that at least one of the plurality of sets of nucleic acid sequences includes an unaligned sequence, wherein the unaligned sequence is a nucleotide sequence that does not align within a threshold number of nucleotides to nucleotides of any known nucleic acid sequences of one or more known microbes predictive of a disease;
determine, for each of the plurality of sets of nucleic acid sequences, whether there is a co-occurrence of (i) a set of nucleic acid sequences of the plurality of sets of nucleic acid sequences including at least the unaligned sequence and (ii) the disease present in a soil sample of the plurality of soil samples associated with the set of nucleic acid sequences;
create a training data set by associating with the disease the unaligned sequence responsive to determining the co-occurrence for a threshold number of the plurality of sets of nucleic acid sequences; and
train the model using the training data set to predict risk of the disease in a test soil sample, wherein the model accounts for a presence of an ameliorative microbe in the test soil sample that reduces the risk of the disease, and wherein the model learns a predictive effect of co-occurrence of the unaligned sequence and the ameliorative microbe on the risk of the disease.
13 . The system of claim 12 , wherein the one or more processors are further configured to:
determine that the unaligned sequence does not correlate to a by-product of the one or more known microbes predictive of the disease.
14 . The system of claim 12 , wherein the one or more processors are further configured to:
train the model with metadata describing a location where the plurality of soil samples is obtained.
15 . The system of claim 12 , wherein the one or more processors are further configured to:
train the model with metadata including one or more of weather patterns, sources of water, fertilizer use, pesticide use, source of seeds, and operational data about a farm from which the plurality of soil samples is sourced.
16 . The system of claim 12 , wherein the one or more processors are further configured to:
determine that the unaligned sequence does not align within a threshold number of nucleotides to the nucleotides of any known nucleic acid sequences by determining absence of a specific loci in the unaligned sequence.
17 . The system of claim 12 , wherein the model is a multi-layered neural network, and wherein the model takes input nucleic acid sequences and outputs phenotypic characteristics.
18 . The system of claim 12 , wherein the disease is citrus greening or strawberry disease.
19 . The system of claim 12 , wherein the one or more processors are further configured to:
determine that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to at least one of the nucleotides of one or more known nucleic acid sequences of the one or more known microbes predictive of the disease; and
determine that presence of the different nucleic acid sequence is predictive of the disease.
20 . The system of claim 12 , wherein the one or more processors are further configured to:
determine that the plurality of sets of nucleic acid sequences includes a different nucleic acid sequence that aligns to nucleotides of nucleic acid sequences of a microbe known to be a suppressor of at least one disease.