IP Library Granted Patent US 12,260,938
Granted Patent B2
US 12,260,938 · App. 17/207,169 · Granted Mar 25, 2025

Machine learning driven gene discovery and gene editing in plants

Inventors: Bradley Zamft (Mountain View, CA); Vikash Singh (Los Angeles, CA); Mathias Voges (Cupertino, CA); Thong Nguyen (Pasadena, CA)
Assignee: HERITABLE AGRICULTURE INC.
G16B40/00G16B5/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,938
App. No.
17/207,169
Granted
Mar 25, 2025
Kind
B2
Abstract

The present disclosure relates to leveraging explainable machine learning methods and feature importance mechanisms as a mechanism for gene discovery and furthermore leveraging the outputs of the gene discovery to recommend ideal gene expression profiles and the requisite genome edits that are conducive to a desired phenotype. Particularly, aspects of the present disclosure are directed to obtaining gene expression profiles for a set of genes measured in a tissue sample of a plant, inputting the gene expression profiles into a prediction model constructed for a task of predicting a phenotype as output data, generating, using the prediction model, the prediction of the phenotype for the plant, analyzing, by an explainable artificial intelligence system, decisions made by the prediction model to predict the phenotype, and identifying a set of candidate gene targets for the phenotype as having a largest contribution or influence on the prediction based on the analyzing.

Claims (85)

1. A method comprising:

obtaining a set of gene expression profiles for a set of genes measured in a tissue sample of a plant;

inputting the set of gene expression profiles into a prediction model having a deep learning architecture constructed for a task of predicting a phenotype as output data by learning relationships or correlations between features of the set of gene expression profiles and the phenotype;

generating, using the prediction model, the prediction of the phenotype for the plant based on the relationships or the correlations between the features of the set of gene expression profiles and the phenotype;

analyzing, by an explainable artificial intelligence system, decisions made by the prediction model to predict the phenotype, wherein the analyzing comprises: (i) generating a set of feature importance scores for the features used in the prediction of the phenotype, wherein the feature importance scores represent estimates of each features contribution or influence on the prediction of the phenotype, and (ii) ranking or otherwise sorting the features based on the feature importance score associated with each of the features, wherein highest ranking or sorted features are identified as having a largest contribution or influence on the prediction of the phenotype;

identifying, based on the ranked or otherwise sorted features, a set of candidate gene targets for the phenotype as having the largest contribution or influence on the prediction of the phenotype;

inputting the set of candidate gene targets into a gene edit modeling system having an architecture constructed to model gene edits and generate ideal gene expression profiles for the phenotype using one or more modeling approaches and the set of candidate gene targets, wherein the one or more modeling approaches generate the ideal gene expression profiles based on the feature importance scores, an optimization algorithm, the prediction model, or any combination thereof;

generating, using the gene edit modeling system, an ideal gene expression profile for the phenotype based on an optimal set of genetic targets for editing of each gene within the set of candidate gene targets, wherein the ideal gene expression profile is a recommendation of gene expression for the set of candidate gene targets for maximizing, minimizing, or otherwise modulating the phenotype; and

making, using a gene editing system, a genetic edit or perturbation to a genome of the plant based on the optimal set of genetic targets for editing each gene and the ideal gene expression profile.

2. The method of claim 1 , wherein the explainable artificial intelligence system uses SHapley Additive explanations, DeepLIFT, integrated gradients, Local Interpretable Model-agnostic Explanations (LIME), Attention-Based Neural Network Models, or Layer-wise Relevance Propagation to analyze the decisions made by the prediction model.

3. The method of claim 1 , wherein the one or more modeling approaches comprises: (i) modeling the gene edits by ascertaining directionality of regulation directly from the feature importance scores, (ii) modeling the gene edits using a Bayesian optimization algorithm, (iii) modeling the gene edits by performing adversarial attack on the prediction model, or (iv) any combination thereof.

4. The method of claim 1 , wherein:

the explainable artificial intelligence system uses SHapley Additive explanations, which generates (i) a set of Shapley values as the feature importance scores for the features used in the prediction of the phenotype and (ii) correlations between the features and the phenotype;

the Shapley values represent estimates of each feature's importance as well as a direction; and

the gene edit modeling system models the gene edits by ascertaining directionality of regulation directly from the Shapley values, wherein the gene edit modeling system is configured to determine how candidate gene targets impact the phenotype through upregulation or downregulation based on the correlations.

5. The method of claim 1 , wherein:

the prediction model is a Gaussian process model; and

the gene edit modeling system models the gene edits using a Bayesian optimization algorithm comprising two components: (i) the Gaussian process model of an underlying Gaussian process function, and (ii) an acquisition function for sampling various data points, wherein the Bayesian optimization algorithm is built based on a sequential search framework that incorporates both exploration and exploitation to direct a search to find a minimum or maximum of an objective function, and wherein the gene edit modeling system is configured to estimate a distribution of the underlying Gaussian process function at the features used in the prediction of the phenotype and use the distribution to direct the sampling by the acquisition function.

6. The method of claim 1 , wherein:

the prediction model is a deep neural network; and

the gene edit modeling system models the gene edits by performing an adversarial attack on the deep neural network, the adversarial attack comprising freezing weights of the deep neural network, and optimizing over a space of constrained inputs to maximize or minimize the phenotype, wherein the optimizing comprises:

identifying the candidate gene targets;

defining an optimization problem for the deep neural network based on an optimal expression corresponding to a set of the candidate gene targets to maximize the phenotype;

defining constraints for gene expression based on the optimal expression; and

generating the optimal set of genetic targets by the adversarial attack.

7. The method of claim 1 , further comprising:

comparing the ideal gene expression profile to a naturally occurring distribution of gene expression for the plant; and

determining a gene edit recommendation for upregulating or downregulating a particular gene, a sub group of genes, or each gene within the ideal gene expression profiles based on the comparing.

8. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform actions including:

obtaining a set of gene expression profiles for a set of genes measured in a tissue sample of a plant;

inputting the set of gene expression profiles into a prediction model having a deep learning architecture constructed for a task of predicting a phenotype as output data by learning relationships or correlations between features of the set of gene expression profiles and the phenotype;

generating, using the prediction model, the prediction of the phenotype for the plant based on the relationships or the correlations between the features of the set of gene expression profiles and the phenotype;

analyzing, by an explainable artificial intelligence system, decisions made by the prediction model to predict the phenotype, wherein the analyzing comprises: (i) generating a set of feature importance scores for the features used in the prediction of the phenotype, wherein the feature importance scores represent estimates of each features contribution or influence on the prediction of the phenotype, and (ii) ranking or otherwise sorting the features based on the feature importance score associated with each of the features, wherein highest ranking or sorted features are identified as having a largest contribution or influence on the prediction of the phenotype;

identifying, based on the ranked or otherwise sorted features, a set of candidate gene targets for the phenotype as having the largest contribution or influence on the prediction of the phenotype;

inputting the set of candidate gene targets into a gene edit modeling system having an architecture constructed to model gene edits and generate ideal gene expression profiles for the phenotype using one or more modeling approaches and the set of candidate gene targets, wherein the one or more modeling approaches generate the ideal gene expression profiles based on the feature importance scores, an optimization algorithm, the prediction model, or any combination thereof;

generating, using the gene edit modeling system, an ideal gene expression profile for the phenotype based on an optimal set of genetic targets for editing of each gene within the set of candidate gene targets, wherein the ideal gene expression profile is a recommendation of gene expression for the set of candidate gene targets for maximizing, minimizing, or otherwise modulating the phenotype; and

instruction to use a gene editing system to make a genetic edit or perturbation to a genome of the plant based on the optimal set of genetic targets for editing each gene and the ideal gene expression profile.

9. The computer-program product of claim 8 , wherein the explainable artificial intelligence system uses SHapley Additive explanations, DeepLIFT, integrated gradients, Local Interpretable Model-agnostic Explanations (LIME), Attention-Based Neural Network Models, or Layer-wise Relevance Propagation to analyze the decisions made by the prediction model.

10. The computer-program product of claim 8 , wherein the one or more modeling approaches comprises: (i) modeling the gene edits by ascertaining directionality of regulation directly from the feature importance scores, (ii) modeling the gene edits using a Bayesian optimization algorithm, (iii) modeling the gene edits by performing adversarial attack on the prediction model, or (iv) any combination thereof.

11. The computer-program product of claim 8 , wherein:

the explainable artificial intelligence system uses SHapley Additive explanations, which generates (i) a set of Shapley values as the feature importance scores for the features used in the prediction of the phenotype and (ii) correlations between the features and the phenotype;

the Shapley values represent estimates of each feature's importance as well as a direction; and

the gene edit modeling system models the gene edits by ascertaining directionality of regulation directly from the Shapley values, wherein the gene edit modeling system is configured to determine how candidate gene targets impact the phenotype through upregulation or downregulation based on the correlations.

12. The computer-program product of claim 8 , wherein:

the prediction model is a Gaussian process model; and

the gene edit modeling system models the gene edits using a Bayesian optimization algorithm comprising two components: (i) the Gaussian process model of an underlying Gaussian process function, and (ii) an acquisition function for sampling various data points, wherein the Bayesian optimization algorithm is built based on a sequential search framework that incorporates both exploration and exploitation to direct a search to find a minimum or maximum of an objective function, and wherein the gene edit modeling system is configured to estimate a distribution of the underlying Gaussian process function at the features used in the prediction of the phenotype and use the distribution to direct the sampling by the acquisition function.

13. The computer-program product of claim 8 , wherein:

the prediction model is a deep neural network; and

the gene edit modeling system models the gene edits by performing an adversarial attack on the deep neural network, the adversarial attack comprising freezing weights of the deep neural network, and optimizing over a space of constrained inputs to maximize or minimize the phenotype, wherein the optimizing comprises:

identifying the candidate gene targets;

defining an optimization problem for the deep neural network based on an optimal expression corresponding to a set of the candidate gene targets to maximize the phenotype;

defining constraints for gene expression based on the optimal expression; and

generating the optimal set of genetic targets by the adversarial attack.

14. The computer-program product of claim 8 , wherein the actions further comprise:

comparing the ideal gene expression profile to a naturally occurring distribution of gene expression for the plant; and

determining a gene edit recommendation for upregulating or downregulating a particular gene, a sub group of genes, or each gene within the ideal gene expression profiles based on the comparing.

15. A system comprising:

one or more data processors; and

a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including:

obtaining a set of gene expression profiles for a set of genes measured in a tissue sample of a plant;

inputting the set of gene expression profiles into a prediction model having a deep learning architecture constructed for a task of predicting a phenotype as output data by learning relationships or correlations between features of the set of gene expression profiles and the phenotype;

generating, using the prediction model, the prediction of the phenotype for the plant based on the relationships or the correlations between the features of the set of gene expression profiles and the phenotype;

analyzing, by an explainable artificial intelligence system, decisions made by the prediction model to predict the phenotype, wherein the analyzing comprises: (i) generating a set of feature importance scores for the features used in the prediction of the phenotype, wherein the feature importance scores represent estimates of each features contribution or influence on the prediction of the phenotype, and (ii) ranking or otherwise sorting the features based on the feature importance score associated with each of the features, wherein highest ranking or sorted features are identified as having a largest contribution or influence on the prediction of the phenotype;

identifying, based on the ranked or otherwise sorted features, a set of candidate gene targets for the phenotype as having the largest contribution or influence on the prediction of the phenotype;

inputting the set of candidate gene targets into a gene edit modeling system having an architecture constructed to model gene edits and generate ideal gene expression profiles for the phenotype using one or more modeling approaches and the set of candidate gene targets, wherein the one or more modeling approaches generate the ideal gene expression profiles based on the feature importance scores, an optimization algorithm, the prediction model, or any combination thereof;

generating, using the gene edit modeling system, an ideal gene expression profile for the phenotype based on an optimal set of genetic targets for editing of each gene within the set of candidate gene targets, wherein the ideal gene expression profile is a recommendation of gene expression for the set of candidate gene targets for maximizing, minimizing, or otherwise modulating the phenotype; and

instruction to use a gene editing system to make a genetic edit or perturbation to a genome of the plant based on the optimal set of genetic targets for editing each gene and the ideal gene expression profile.

16. The system of claim 15 , wherein the one or more modeling approaches comprises: (i) modeling the gene edits by ascertaining directionality of regulation directly from the feature importance scores, (ii) modeling the gene edits using a Bayesian optimization algorithm, (iii) modeling the gene edits by performing adversarial attack on the prediction model, or (iv) any combination thereof.

17. The system of claim 15 , wherein:

the explainable artificial intelligence system uses SHapley Additive explanations, which generates (i) a set of Shapley values as the feature importance scores for the features used in the prediction of the phenotype and (ii) correlations between the features and the phenotype;

the Shapley values represent estimates of each feature's importance as well as a direction; and

the gene edit modeling system models the gene edits by ascertaining directionality of regulation directly from the Shapley values, wherein the gene edit modeling system is configured to determine how candidate gene targets impact the phenotype through upregulation or downregulation based on the correlations.

18. The system of claim 15 , wherein:

the prediction model is a Gaussian process model; and

the gene edit modeling system models the gene edits using a Bayesian optimization algorithm comprising two components: (i) the Gaussian process model of an underlying Gaussian process function, and (ii) an acquisition function for sampling various data points, wherein the Bayesian optimization algorithm is built based on a sequential search framework that incorporates both exploration and exploitation to direct a search to find a minimum or maximum of an objective function, and wherein the gene edit modeling system is configured to estimate a distribution of the underlying Gaussian process function at the features used in the prediction of the phenotype and use the distribution to direct the sampling by the acquisition function.

19. The system of claim 15 , wherein:

the prediction model is a deep neural network; and

the gene edit modeling system models the gene edits by performing an adversarial attack on the deep neural network, the adversarial attack comprising freezing weights of the deep neural network, and optimizing over a space of constrained inputs to maximize or minimize the phenotype, wherein the optimizing comprises:

identifying the candidate gene targets;

defining an optimization problem for the deep neural network based on an optimal expression corresponding to a set of the candidate gene targets to maximize the phenotype;

defining constraints for gene expression based on the optimal expression; and

generating the optimal set of genetic targets by the adversarial attack.

20. The system of claim 15 , wherein the actions further comprise:

comparing the ideal gene expression profile to a naturally occurring distribution of gene expression for the plant; and

determining a gene edit recommendation for upregulating or downregulating a particular gene, a sub group of genes, or each gene within the ideal gene expression profiles based on the comparing.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: X DEVELOPMENT LLC
To: HERITABLE AGRICULTURE INC.
Reel/Frame 069767/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2021
From: ZAMFT, BRADLEY; SINGH, VIKASH; VOGES, MATHIAS; NGUYEN, THONG
To: X DEVELOPMENT LLC
Reel/Frame 055793/0805 →
Continuity (1)
Related Publication 20220301658A1 · Sep 22, 2022
References Cited (18)
US 11618896B2 · Zhang · 2023 [cited by examiner]
US 20140220568A1 · Inze et al. · 2014 [cited by applicant]
US 20180285520A1 · Butruille et al. · 2018 [cited by applicant]
US 20200202241A1 · Abeliuk · 2020 [cited by examiner]
WO 2021035164A1 · 2021 [cited by applicant]
Wu et al., “Scalable Planning with Deep Neural Network Learned Transition Models”, Journal of Artificial Intelligence Research, Jul. 2020, vol. 68, https://doi.org/10.1613/jair.1.11829, 571-606 (Year: 2020). [cited by examiner]
Arsham Ghahramani, Fiona M. Watt, Nicholas M. Luscombe; “Generative adversarial networks simulate gene expression and predict perturbations in single cells”; bioRxiv 262501; Jul. 30, 2018; doi: https://doi.org/10.1101/2… [cited by examiner]
Lundberg, Scott M., and Su-In Lee. “A unified approach to interpreting model predictions.” Advances in neural information processing systems 30 (2017). (Year: 2017). [cited by examiner]
Borrill et al., “Identification of Transcription Factors Regulating Senescence in Wheat Through Gene Regulatory Network Modelling”, Plant Physiololgy, vol. 180, Jul. 2019, 16 pages. [cited by applicant]
Jin et al., “Auto-Keras: An Efficient Neural Architecture Search System”, Applied Data Science Track Paper, 2019, 11 pages. [cited by applicant]
Kannan et al., “Combining Gene Network, Metablolic and Leaf-Level Models Shows Means to Future-Proof Soybean Photosynthesis Under Rising CO2”, In Silico Plants, vol. 1, Issue 1, 2019, 18 paages, pages. [cited by applicant]
Kawakatsu et al., “Epigenomic Diversity in a Global Collection of [cited by applicant]
Perez-Enciso et al., “A Guide on Deep Learning for Complex Trait Genomic Prediction”, Genes, 10, 553,, Jul. 2019, 19 pages. [cited by applicant]
Vats et al., “Genome Editing in Plants: Exploration of Technological Advancements and Challenges”, Cells, 8 (11), 1386, Nov. 4, 2019, 39 pages. [cited by applicant]
Wang et al., “Deep Learning for Plant Genomics and Crop Improvement”, Current Opinion in Plant Biology, retrieved from www. science direct.com, 2020, 8 pages. [cited by applicant]
Harfouche et al., “Accelerating Climate Resilient Plant Breeding by Applying Next-Generation Artificial Intelligence”, Trends in Biotechnology, vol. 37, No. 11, Jun. 21, 2019, pp. 1217-1235. [cited by applicant]
International Application No. PCT/US2021/060694, International Search Report and Written Opinion, Mailed On Mar. 7, 2022, 15 pages. [cited by applicant]
Mlone et al., “Explainable Artificial Intelligence: a Systematic Review”, arxiv.org, Available Online at URL: https://arxiv.org/abs/2006.00093, May 2020, 81 pages. [cited by applicant]