IP Library › Granted Patent US 12,738,081
Granted Patent B2
US 12,738,081 · App. 18/779,917 · Granted Sep 15, 2026

Utilizing machine learning and digital embedding processes to generate digital maps of biology and user interfaces for evaluating map efficacy

Inventors: Nathan Henry Lazar (Salt Lake City, UT); Conor Austin Forsman Tillinghast (Salt Lake City, UT); James Douglas Jensen (Farmington, UT); James Benjamin Taylor (Midlothian, VA); Berton Allen Earnshaw (Cedar Hills, UT); Marta Marie Fay (Salt Lake City, UT); Renat Nailevich Khaliullin (Salt Lake City, UT); Jacob Carter Cooper (Sandy, UT); Imran Saeedul Haque (Salt Lake City, UT); Seyhmus Guler (Salt Lake City, UT); Kyle Rollins Hansen (Kaysville, UT); Safiye Celik (Sudbury, MA)
Assignee: Recursion Pharmaceuticals, Inc.
G06V20/698G06F16/51G06F16/583G06T7/0012G06T7/35G06V10/761G06V10/82G06V10/95G16B40/00G16B40/20G16B50/30G06T2200/24G06T2207/20081G06T2207/20084G06T2207/30072
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,081
App. No.
18/779,917
Granted
Sep 15, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing machine learning and digital embedding processes to generate digital maps of biology and user interfaces for evaluating map efficacy. In particular, in one or more embodiments, the disclosed systems receive perturbation data for a plurality of perturbation experiment units corresponding to a plurality of perturbation classes. Further, the systems generate, utilizing a machine learning model, a plurality of perturbation experiment unit embeddings from the perturbation data. Additionally, the systems align, utilizing an alignment model, the plurality of perturbation experiment unit embeddings to generate aligned perturbation unit embeddings. Moreover, the systems aggregate the aligned perturbation unit embeddings to generate aggregated embeddings. Furthermore, the systems generate perturbation comparisons utilizing the perturbation-level embeddings.

Claims (49)

1 . A computer-implemented method comprising:

receiving, from a client device, a similarity request for two or more cell perturbations;

generating similarity measures between the two or more cell perturbations by comparing aggregated embeddings corresponding to the two or more cell perturbations, wherein the aggregated embeddings are aggregated from perturbation experiment unit embeddings generated utilizing a machine learning model from perturbation data; and

in response to the similarity request, providing, for display via the client device, a perturbation similarity map reflecting the similarity measures between the two or more cell perturbations.

2 . The computer-implemented method of claim 1 , wherein receiving, from the client device, the similarity request for two or more cell perturbations comprises:

receiving a compound-gene similarity request between a compound perturbation and a gene perturbation; or

receiving a compound-compound similarity request between a first compound perturbation and a second compound perturbation.

3 . The computer-implemented method of claim 1 , wherein generating the similarity measures comprises comparing a first aggregated embedding for a first perturbation and a second aggregated embedding for a second perturbation in a machine learning feature space to generate a similarity measure.

4 . The computer-implemented method of claim 1 , further comprising generating the perturbation experiment unit embeddings utilizing the machine learning model from the perturbation data by:

capturing phenomic images of cells exposed to perturbations; and

generating, utilizing the machine learning model, perturbation experiment unit embeddings from the phenomic images of the cells.

5 . The computer-implemented method of claim 1 , further comprising generating the perturbation experiment unit embeddings utilizing the machine learning model from the perturbation data by:

generating transcriptomic profiles of cells exposed to perturbations; and

generating the perturbation experiment unit embeddings from the transcriptomic profiles.

6 . The computer-implemented method of claim 1 , wherein providing the perturbation similarity map reflecting the similarity measures between the two or more cell perturbations for display comprises providing, for display via the client device, a heatmap comprising cells having shading that represents the similarity measures.

7 . The computer-implemented method of claim 1 , wherein providing the perturbation similarity map further comprises providing, for display via the client device, a confidence measure for one or more of the aggregated embeddings.

8 . A system comprising:

at least one processor; and

at least one non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, cause the system to:

receive, from a client device, a similarity request for two or more cell perturbations;

generate similarity measures between the two or more cell perturbations by comparing aggregated embeddings corresponding to the two or more cell perturbations, wherein the aggregated embeddings are aggregated from perturbation experiment unit embeddings generated utilizing a machine learning model from perturbation data; and

in response to the similarity request, provide, for display via the client device, a perturbation similarity map reflecting the similarity measures between the two or more cell perturbations.

9 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to:

receive a compound-gene similarity request between a compound perturbation and a gene perturbation;

receive a compound-compound similarity request between a first compound perturbation and a second compound perturbation; and

provide, for display via the client device, the perturbation similarity map reflecting a first similarity measure between the compound perturbation and the gene perturbation and a second similarity measure between the first compound perturbation and the second compound perturbation.

10 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the similarity measures by comparing a first aggregated embedding for a first perturbation and a second aggregated embedding for a second perturbation in a machine learning feature space to generate a similarity measure.

11 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the perturbation experiment unit embeddings utilizing the machine learning model from the perturbation data by:

capturing phenomic images of cells exposed to perturbations; and

generating, utilizing the machine learning model, perturbation experiment unit embeddings from the phenomic images of the cells.

12 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the perturbation experiment unit embeddings utilizing the machine learning model from the perturbation data by:

generating transcriptomic profiles of cells exposed to perturbations; and

generating the perturbation experiment unit embeddings from the transcriptomic profiles.

13 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to provide the perturbation similarity map reflecting the similarity measures between the two or more cell perturbations for display by providing, for display via the client device, a heatmap comprising cells having shading that represents the similarity measures.

14 . The system of claim 8 , further comprising instructions that, when executed by the at least one processor, cause the system to provide the perturbation similarity map by providing, for display via the client device, a confidence measure for one or more of the aggregated embeddings.

15 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:

receive, from a client device, a similarity request for two or more cell perturbations;

generate similarity measures between the two or more cell perturbations by comparing aggregated embeddings corresponding to the two or more cell perturbations, wherein the aggregated embeddings are aggregated from perturbation experiment unit embeddings generated utilizing a machine learning model from perturbation data; and

in response to the similarity request, provide, for display via the client device, a perturbation similarity map reflecting the similarity measures between the two or more cell perturbations.

16 . The non-transitory computer-readable medium of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

receive a compound-gene similarity request between a compound perturbation and a gene perturbation;

receive a compound-compound similarity request between a first compound perturbation and a second compound perturbation; and

provide, for display via the client device, the perturbation similarity map reflecting a first similarity measure between the compound perturbation and the gene perturbation and a second similarity measure between the first compound perturbation and the second compound perturbation.

17 . The non-transitory computer-readable medium of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the similarity measures by comparing a first aggregated embedding for a first perturbation and a second aggregated embedding for a second perturbation in a machine learning feature space to generate a similarity measure.

18 . The non-transitory computer-readable medium of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the perturbation experiment unit embeddings utilizing the machine learning model from the perturbation data by:

capturing phenomic images of cells exposed to perturbations; and

generating, utilizing the machine learning model, perturbation experiment unit embeddings from the phenomic images of the cells.

19 . The non-transitory computer-readable medium of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the computing device to provide the perturbation similarity map reflecting the similarity measures between the two or more cell perturbations for display by providing, for display via the client device, a heatmap comprising cells having shading that represents the similarity measures.

20 . The non-transitory computer-readable medium of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the computing device to provide the perturbation similarity map by providing, for display via the client device, a confidence measure for one or more of the aggregated embeddings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2024
From: LAZAR, NATHAN HENRY; TILLINGHAST, CONOR AUSTIN FORSMAN; JENSEN, JAMES DOUGLAS; TAYLOR, JAMES BENJAMIN; EARNSHAW, BERTON ALLEN; FAY, MARTA MARIE; KHALIULLIN, RENAT NAILEVICH; COOPER, JACOB CARTER; HAQUE, IMRAN SAEEDUL; GULER, SEYHMUS; HANSEN, KYLE ROLLINS; CELIK, SAFIYE
To: RECURSION PHARMACEUTICALS, INC.
Reel/Frame 068138/0548 →
Continuity (3)
Continuation 18393041 · Dec 21, 2023
Provisional Application 63582702 · Sep 14, 2023
Related Publication 20250095146A1 · Mar 20, 2025
References Cited (57)
US 7047410B1 · Shin · 2006 [cited by applicant]
US 8812526B2 · Ramer et al. · 2014 [cited by applicant]
US 10146914B1 · Victors et al. · 2018 [cited by applicant]
US 10769501B1 · Ando · 2020 [cited by examiner]
US 12073638B1 · Lazar et al. · 2024 [cited by applicant]
US 12079992B1 · Lazar et al. · 2024 [cited by applicant]
US 12119090B1 · Kraus et al. · 2024 [cited by applicant]
US 12373950B1 · Minasian et al. · 2025 [cited by applicant]
US 12374429B1 · Fay et al. · 2025 [cited by applicant]
US 20100223276A1 · Al-Shameri et al. · 2010 [cited by applicant]
US 20150320319A1 · Alfano et al. · 2015 [cited by applicant]
US 20190114390A1 · Donner · 2019 [cited by examiner]
US 20200126637A1 · Xu et al. · 2020 [cited by applicant]
US 20210133976A1 · Carmi · 2021 [cited by applicant]
US 20210257044A1 · Califano et al. · 2021 [cited by applicant]
US 20210366577A1 · Koller et al. · 2021 [cited by applicant]
US 20210374553A1 · Li et al. · 2021 [cited by applicant]
US 20220180975A1 · Regev et al. · 2022 [cited by applicant]
US 20230416307A1 · Lopez et al. · 2023 [cited by applicant]
US 20240112447A1 · Sawada et al. · 2024 [cited by applicant]
US 20240304009A1 · Yoon et al. · 2024 [cited by applicant]
US 20260051164A1 · Nykl et al. · 2026 [cited by applicant]
U.S. Appl. No. 18/526,729, May 27, 2025, Notice of Allowance. [cited by applicant]
Yihan Zhang, Luing Yang, and Vladimir Brusic, “Automation of Gene Expression Profile Analysis in Single Cell Data”, 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Seoul, Korea (South), 2020… [cited by applicant]
U.S. Appl. No. 18/526,707, Jan. 15, 2026, Office Action. [cited by applicant]
U.S. Appl. No. 18/526,742, Apr. 3, 2026, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/526,729, Jan. 16, 2025, Office Action. [cited by applicant]
Akshay Agrawal, Alnur Ali, Stephen Boyd, et al. Minimum-distortion embedding. Foundations and Trends® in Machine Learning, 14(3):211-378, 2021. [cited by applicant]
Arthur Liberzon, Chet Birger, Helga Thorvaldsdóttir, Mahmoud Ghandi, Jill P Mesirov, and Pablo Tamayo. The molecular signatures database hallmark gene set collection. Cell systems, 1(6):417-425, 2015. [cited by applicant]
Atray Dixit, Oren Parnas, Biyu Li, Jenny Chen, Charles P Fulco, Livnat Jerby-Arnon, Nemanja D Marjanovic, Danielle Dionne, Tyler Burks, Raktima Raychowdhury, et al. Perturb-seq: dissecting molecular circuits with scalab… [cited by applicant]
Aurora S Blucher, Safiye Celik, James D Jensen, James Taylor, Michael F Cuccarese, Jacob C Cooper, Jacob M Rinaldi, Carl Brooks, Michael A Statnick, Marta Fay, Nathan Lazar, Berton Earnshaw, and Imran S Haque. Poster: M… [cited by applicant]
D Michael Ando, Cory Y McLean, and Marc Berndl. Improving phenotypic measurements in high-content imaging screens. BioRxiv, p. 161422, 2017. [cited by applicant]
Gabor J Szekely. Potential and kinetic energy in statistics. Lecture Notes, Budapest Institute, 1989. [cited by applicant]
Gökcen Eraslan, Lukas M Simon, Maria Mircea, Nikola S Mueller, and Fabian J Theis. Singlecell rna-seq denoising using a deep count autoencoder. Nature communications, 10(1):1-14, 2019. [cited by applicant]
John W Tukey. Mathematics and the picturing of data. In Proceedings of the International Congress of Mathematicians, Vancouver, 1975, vol. 2, pp. 523-531, 1975. [cited by applicant]
Joseph M Replogle, Reuben A Saunders, Angela N Pogson, Jeffrey A Hussmann, Alexander Lenail, Alina Guna, Lauren Mascibroda, Eric J Wagner, Karen Adelman, Gila Lithwick-Yanai, et al. Mapping information-rich genotype-phe… [cited by applicant]
Kevin Drew, John B Wallingford, and Edward M Marcotte. hu.map 2.0: integration of over 15,000 proteomic experiments builds a global compendium of human multiprotein assemblies. Mol Syst Biol, 17(5):e10016, 2021. [cited by applicant]
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. Advances in neural information processing systems, 28, 2015. [cited by applicant]
Krzysztof Polanski, Matthew D Young, Zhichao Miao, Kerstin B Meyer, Sarah A Teichmann, and Jong-Eun Park. Bbknn: fast batch alignment of single cell transcriptomes. Bioinformatics, 36(3):964-965, 2020. [cited by applicant]
Laleh Haghverdi, Aaron TL Lun, Michael D Morgan, and John C Marioni. Batch effects in single-cell rna-sequencing data are corrected by matching mutual nearest neighbors. Nature biotechnology, 36(5):421-427, 2018. [cited by applicant]
Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. [cited by applicant]
Luana Licata, Prisca Lo Surdo, Marta Iannuccelli, Alessandro Palma, Elisa Micarelli, Livia Perfetto, Daniele Peluso, Alberto Calderone, Luisa Castagnoli, and Gianni Cesareni. Signor 2.0, the signaling network open resou… [cited by applicant]
Madalina Giurgiu, Julian Reinhard, Barbara Brauner, Irmtraud Dunger-Kaltenbach, Gisela Fobo, Goar Frishman, Corinna Montrone, and Andreas Ruepp. Corum: the comprehensive resource of mammalian protein complexes—2019. Nuc… [cited by applicant]
Marc Gillespie, Bijay Jassal, Ralf Stephan, Marija Milacic, Karen Rothfels, Andrea Senff-Ribeiro, Johannes Griss, Cristoffer Sevilla, Lisa Matthews, Chuqiao Gong, et al. The reactome pathway knowledgebase 2022. Nucleic … [cited by applicant]
Maria L Rizzo and Gábor J Székely. Energy distance. wiley interdisciplinary reviews: Computational statistics, 8(1):27-38, 2016. [cited by applicant]
Mark-Anthony Bray, Shantanu Singh, Han Han, Chadwick T Davis, Blake Borgeson, Cathy Hartland, Maria Kost-Alimova, Sigrun M Gustafsdottir, Christopher C Gibson, and Anne E Carpenter. Cell painting, a high-content image-b… [cited by applicant]
Michael F Cuccarese, Berton A Earnshaw, Katie Heiser, Ben Fogelson, Chadwick T Davis, Peter F McLean, Hannah B Gordon, Kathleen-Rose Skelly, Fiona L Weathersby, Vlad Rodic, Ian K Quigley, Elissa D Pastuzyn, Brandon M Me… [cited by applicant]
Nathan Lazar, et al. High-Resolution Genome-wide Mapping of Chromosome-arm-scale Truncations Induced by CRISPR-Cas9 Editing published in bioRxiv on Apr. 15, 2023. [cited by applicant]
Romain Lopez, Jeffrey Regier, Michael B Cole, Michael I Jordan, and Nir Yosef. Deep generative modeling for single-cell transcriptomics. Nature methods, 15(12):1053-1058, 2018. [cited by applicant]
Sivanandan et al. “A Pooled Cell Painting CRISPR Screening Platform Enables de novo Inference of Gene Function by Self-supervised Deep Learning”, https://www.biorxiv.org/contenU10.1101/2023.08.13.553051v3.abstract, bioR… [cited by applicant]
U.S. Appl. No. 18/393,041, Feb. 23, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/393,041, May 3, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/392,989, Mar. 1, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/392,989, Apr. 29, 2024, Notice of Allowance. [cited by applicant]
Robin M Meyers, et al. “Computational correction of copy number effect improves specificity of CRISPR-Cas9 essentiality screens in cancer cells.” Nature genetics 49.12 (2017): pp. 1779-1788) (Year: 2017). [cited by applicant]
U.S. Appl. No. 18/526,707, Apr. 14, 2026, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/778,647, May 14, 2026, Office Action. [cited by applicant]