IP Library › Granted Patent US 12,374,429
Granted Patent B1
US 12,374,429 · App. 18/526,729 · Granted Jul 29, 2025

Utilizing machine learning models to synthesize perturbation data to generate perturbation heatmap graphical user interfaces

Inventors: Marta Marie Fay (Salt Lake City, UT); August Orvis Allen (Boulder, CO); Eugene Yin-Chung Ting (Toronto, CA); Lina Maria Nilsson (Salt Lake City, UT); Condie Thomas Swallow, II (West Valley City, UT); Michael Haines (Salt Lake City, UT); Denton Hallar Greenfield (Evansville, IN); Kristin Ann Clark (Lehi, UT); Lovina Roundy (Orem, UT); Michael Joseph Uloth (Dundas, CA); Sara Marjean Moore (Boise, ID); Shweta Deepchand Bhandare (Boulder, CO); Ted Douglas Monchamp (Nashua, NH); Summer Walid Elias (Salt Lake City, UT); Berton Allen Earnshaw (Cedar Hills, UT); Mason Lemoyne Victors (Riverton, UT); Safiye Celik (Sudbury, MA); James Benjamin Taylor (Midlothian, VA); Andrew David Blevins (Salt Lake City, UT); James Douglas Jensen (Farmington, UT); Jacob Carter Cooper (Sandy, UT); Conor Austin Forsman Tillinghast (Salt Lake City, UT); Seyhmus Guler (Salt Lake City, UT); Kyle Rollins Hansen (Kaysville, UT); Sarah Jordan DeVore (Salt Lake City, UT); Tongzhou Shen (Surrey, CA)
Assignee: Recursion Pharmaceuticals, Inc.
G16B50/30G06F16/51G06F16/583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,429
App. No.
18/526,729
Granted
Jul 29, 2025
Kind
B1
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for embedding perturbation data via a machine learning model and filtering, aligning, and aggregating the embeddings to generate a genome-wide perturbation database for real-time generation of perturbation heatmaps. In particular, in one or more embodiments, the disclosed systems can receive a plurality of perturbation images portraying cells from a plurality of wells corresponding to a plurality of cell perturbations. Further, the systems can generate, utilizing a machine learning model, a plurality of well-level image embeddings from the plurality of perturbation images. Moreover, the systems can align, utilizing an alignment model, the plurality of well-level image embeddings to generate aligned well-level image embeddings. Additionally, the systems can aggregate, according to perturbations of one or more perturbation experiments, the well-level image embeddings to generate perturbation-level image embeddings. Furthermore, the systems can generate perturbation comparisons utilizing the perturbation-level image embeddings.

Claims (83)

1. A computer-implemented method comprising:

identifying a plurality of initial cell image embeddings;

aggregating the plurality of initial cell image embeddings according to cell perturbations to generate cell image embeddings;

generating a dataframe comprising the cell image embeddings corresponding to the cell perturbations;

receiving, from a client device, a similarity query comprising a plurality of cell perturbations;

in response to receiving the similarity query, accessing cell image embeddings corresponding to the plurality of cell perturbations from the dataframe;

determining similarity measures for the plurality of cell perturbations from the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query; and

transmitting, to the client device, a query response comprising the similarity measures from the cell image embeddings.

2. The computer-implemented method of claim 1 , further comprising:

generating an additional dataframe comprising aggregated embedding metadata of the cell image embeddings of the dataframe;

identifying an additional plurality of initial cell image embeddings incorporated into a perturbation database and metadata of the additional plurality of initial cell image embeddings; and

regenerating the dataframe to include the additional plurality of initial cell image embeddings and the additional dataframe to include the metadata of the additional plurality of initial cell image embeddings.

3. The computer-implemented method of claim 1 , further comprising:

aggregating metadata corresponding to the plurality of initial cell image embeddings to generate aggregated embedding metadata;

generating an additional dataframe comprising the aggregated embedding metadata of the cell image embeddings of the dataframe;

accessing the cell image embeddings from the dataframe and corresponding aggregated embedding metadata from the additional dataframe; and

transmitting the query response by transmitting the cell image embeddings and the aggregated embedding metadata to the client device.

4. The computer-implemented method of claim 3 , wherein determining the similarity measures for the plurality of cell perturbations from the cell image embeddings comprises at least one of:

comparing a first cell image embedding of the cell image embeddings with a second cell image embedding of the cell image embeddings; or

comparing the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query with additional cell image embeddings corresponding to additional cell perturbations.

5. The computer-implemented method of claim 4 , further comprising:

receiving additional cell image embeddings; and

aggregating the additional cell image embeddings and the plurality of initial cell image embeddings to generate updated cell image embeddings.

6. The computer-implemented method of claim 5 , further comprising:

receiving, from an additional client device, an additional similarity query comprising an additional set of cell perturbations;

accessing the updated cell image embeddings corresponding to the additional set of cell perturbations; and

determining additional similarity measures for the additional set of cell perturbations from the updated cell image embeddings corresponding to the additional set of cell perturbations of the additional similarity query.

7. The computer-implemented method of claim 1 , wherein determining the similarity measures between the plurality of cell perturbations by comparing the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query comprises determining a cosine similarity between the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query.

8. The computer-implemented method of claim 1 , wherein receiving, from the client device, the similarity query comprising the plurality of cell perturbations comprises receiving at least one of a gene perturbation query or a compound perturbation query.

9. The computer-implemented method of claim 1 , wherein transmitting, to the client device, the query response comprising the similarity measures comprises transmitting, to the client device, a perturbation heatmap portraying the similarity measures.

10. A system comprising:

at least one processor; and

at least one non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, cause the system to:

identify a plurality of initial cell image embeddings;

aggregate the plurality of initial cell image embeddings according to cell perturbations to generate cell image embeddings;

generate a dataframe comprising the cell image embeddings corresponding to the cell perturbations;

receive, from a client device, a similarity query comprising a plurality of cell perturbations;

in response to receiving the similarity query, access cell image embeddings corresponding to the plurality of cell perturbations from the dataframe;

determine similarity measures for the plurality of cell perturbations from the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query; and

transmit, to the client device, a query response comprising the similarity measures from the cell image embeddings.

11. The system of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the system to:

generate an additional dataframe comprising aggregated embedding metadata of the cell image embeddings of the dataframe;

identify an additional plurality of initial cell image embeddings incorporated into a perturbation database and metadata of the additional plurality of initial cell image embeddings; and

regenerate the dataframe to include the additional plurality of initial cell image embeddings and the additional dataframe to include the metadata of the additional plurality of initial cell image embeddings.

12. The system of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the system to:

aggregate metadata corresponding to the plurality of initial cell image embeddings to generate aggregated embedding metadata;

generate an additional dataframe comprising the aggregated embedding metadata of the cell image embeddings of the dataframe;

access the cell image embeddings from the dataframe and corresponding aggregated embedding metadata from the additional dataframe; and

transmit the query response by transmitting the cell image embeddings and the aggregated embedding metadata to the client device.

13. The system of claim 12 , further comprising instructions that, when executed by the at least one processor, cause the system to determine the similarity measures for the plurality of cell perturbations from the cell image embeddings by performing at least one of:

comparing a first cell image embedding of the cell image embeddings with a second cell image embedding of the cell image embeddings; or

comparing the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query with additional cell image embeddings corresponding to additional cell perturbations.

14. The system of claim 13 , further comprising instructions that, when executed by the at least one processor, cause the system to:

receive additional cell image embeddings; and

aggregate the additional cell image embeddings and the plurality of initial cell image embeddings to generate updated cell image embeddings.

15. The system of claim 14 , further comprising instructions that, when executed by the at least one processor, cause the system to:

receive, from an additional client device, an additional similarity query comprising an additional set of cell perturbations;

access the updated cell image embeddings corresponding to the additional set of cell perturbations; and

determine additional similarity measures for the additional set of cell perturbations from the updated cell image embeddings corresponding to the additional set of cell perturbations of the additional similarity query.

16. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause a computing device to:

identify a plurality of initial cell image embeddings;

aggregate the plurality of initial cell image embeddings according to cell perturbations to generate cell image embeddings;

generate a dataframe comprising the cell image embeddings corresponding to the cell perturbations;

receive, from a client device, a similarity query comprising a plurality of cell perturbations;

in response to receiving the similarity query, access cell image embeddings corresponding to the plurality of cell perturbations from the dataframe;

determine similarity measures for the plurality of cell perturbations from the cell image embeddings corresponding to the plurality of cell perturbations of the similarity query; and

transmit, to the client device, a query response comprising the similarity measures from the cell image embeddings.

17. The non-transitory computer-readable storage medium of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

generate an additional dataframe comprising aggregated embedding metadata of the cell image embeddings of the dataframe;

identify an additional plurality of initial cell image embeddings incorporated into a perturbation database and metadata of the additional plurality of initial cell image embeddings; and

regenerate the dataframe to include the additional plurality of initial cell image embeddings and the additional dataframe to include the metadata of the additional plurality of initial cell image embeddings.

18. The non-transitory computer-readable storage medium of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

aggregate metadata corresponding to the plurality of initial cell image embeddings to generate aggregated embedding metadata;

generate an additional dataframe comprising the aggregated embedding metadata of the cell image embeddings of the dataframe;

accessing the cell image embeddings from the dataframe and corresponding aggregated embedding metadata from the additional dataframe; and

transmitting the query response by transmitting the cell image embeddings and the aggregated embedding metadata to the client device.

19. The non-transitory computer-readable storage medium of claim 18 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

receive additional cell image embeddings; and

aggregate the additional cell image embeddings and the plurality of initial cell image embeddings to generate updated cell image embeddings.

20. The non-transitory computer-readable storage medium of claim 19 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

receive, from an additional client device, an additional similarity query comprising an additional set of cell perturbations;

access the updated cell image embeddings corresponding to the additional set of cell perturbations; and

determine additional similarity measures for the additional set of cell perturbations from the updated cell image embeddings corresponding to the additional set of cell perturbations of the additional similarity query.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2023
From: FAY, MARTA MARIE; ALLEN, AUGUST ORVIS; TING, EUGENE YIN-CHUNG; NILSSON, LINA MARIA; SWALLOW, CONDIE THOMAS, II; HAINES, MICHAEL; GREENFIELD, DENTON HALLAR; CLARK, KRISTIN ANN; ULOTH, MICHAEL JOSEPH; MOORE, SARA MARJEAN; BHANDARE, SHWETA DEEPCHAND; MONCHAMP, TED DOUGLAS; ELIAS, SUMMER WALID; EARNSHAW, BERTON ALLEN; VICTORS, MASON LEMOYNE; CELIK, SAFIYE; TAYLOR, JAMES BENJAMIN; BLEVINS, ANDREW DAVID; JENSEN, JAMES DOUGLAS; COOPER, JACOB CARTER; TILLINGHAST, CONOR AUSTIN FORSMAN; GULER, SEYHMUS; HANSEN, KYLE ROLLINS; DEVORE, SARAH JORDAN; SHEN, TONGZHOU
To: RECURSION PHARMACEUTICALS, INC.
Reel/Frame 065738/0218 →
AGREEMENT Recorded Dec 1, 2023
From: ROUNDY, LOVINA
To: RECURSION PHARMACEUTICALS, INC.
Reel/Frame 065744/0694 →
Continuity (1)
Provisional Application 63582702 · Sep 14, 2023
References Cited (41)
US 7047410B1 · Shin · 2006 [cited by examiner]
US 8812526B2 · Ramer · 2014 [cited by examiner]
US 10146914B1 · Victors et al. · 2018 [cited by applicant]
US 10769501B1 · Ando et al. · 2020 [cited by applicant]
US 12119090B1 · Kraus et al. · 2024 [cited by applicant]
US 20100223276A1 · Al-Shameri · 2010 [cited by examiner]
US 20150320319A1 · Alfano · 2015 [cited by examiner]
US 20190114390A1 · Donner et al. · 2019 [cited by applicant]
US 20210133976A1 · Carmi · 2021 [cited by applicant]
US 20210366577A1 · Koller et al. · 2021 [cited by applicant]
US 20210374553A1 · Li · 2021 [cited by examiner]
US 20220180975A1 · Regev et al. · 2022 [cited by applicant]
US 20240112447A1 · Sawada · 2024 [cited by examiner]
US 20240304009A1 · Yoon · 2024 [cited by examiner]
Akshay Agrawal, Alnur Ali, Stephen Boyd, et al. Minimum-distortion embedding. Foundations and Trends® in Machine Learning, 14(3):211-378, 2021. [cited by applicant]
Arthur Liberzon, Chet Birger, Helga Thorvaldsdóttir, Mahmoud Ghandi, Jill P Mesirov, and Pablo Tamayo. The molecular signatures database hallmark gene set collection. Cell systems, 1(6):417-425, 2015. [cited by applicant]
Atray Dixit, Oren Parnas, Biyu Li, Jenny Chen, Charles P Fulco, Livnat Jerby-Arnon, Nemanja D Marjanovic, Danielle Dionne, Tyler Burks, Raktima Raychowdhury, et al. Perturb-seq: dissecting molecular circuits with scalab… [cited by applicant]
Aurora S Blucher, Safiye Celik, James D Jensen, James Taylor, Michael F Cuccarese, Jacob C Cooper, Jacob M Rinaldi, Carl Brooks, Michael A Statnick, Marta Fay, Nathan Lazar, Berton Earnshaw, and Imran S Haque. Poster: M… [cited by applicant]
D Michael Ando, Cory Y McLean, and Marc Berndl. Improving phenotypic measurements in high-content imaging screens. BioRxiv, p. 161422, 2017. [cited by applicant]
Gabor J Szekely. Potential and kinetic energy in statistics. Lecture Notes, Budapest Institute, 1989. [cited by applicant]
Gökcen Eraslan, Lukas M Simon, Maria Mircea, Nikola S Mueller, and Fabian J Theis. Singlecell rna-seq denoising using a deep count autoencoder. Nature communications, 10(1):1-14, 2019. [cited by applicant]
John W Tukey. Mathematics and the picturing of data. In Proceedings of the International Congress of Mathematicians, Vancouver, 1975, vol. 2, pp. 523-531, 1975. [cited by applicant]
Joseph M Replogle, Reuben A Saunders, Angela N Pogson, Jeffrey A Hussmann, Alexander Lenail, Alina Guna, Lauren Mascibroda, Eric J Wagner, Karen Adelman, Gila Lithwick-Yanai, et al. Mapping information-rich genotype-phe… [cited by applicant]
Kevin Drew, John B Wallingford, and Edward M Marcotte. hu.map 2.0: integration of over 15,000 proteomic experiments builds a global compendium of human multiprotein assemblies. Mol Syst Biol, 17(5):e10016, 2021. [cited by applicant]
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. Advances in neural information processing systems, 28, 2015. [cited by applicant]
Krzysztof Polanski, Matthew D Young, Zhichao Miao, Kerstin B Meyer, Sarah A Teichmann, and Jong-Eun Park. Bbknn: fast batch alignment of single cell transcriptomes. Bioinformatics, 36(3):964-965, 2020. [cited by applicant]
Laleh Haghverdi, Aaron TL Lun, Michael D Morgan, and John C Marioni. Batch effects in single-cell rna-sequencing data are corrected by matching mutual nearest neighbors. Nature biotechnology, 36(5):421-427, 2018. [cited by applicant]
Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. [cited by applicant]
Luana Licata, Prisca Lo Surdo, Marta Iannuccelli, Alessandro Palma, Elisa Micarelli, Livia Perfetto, Daniele Peluso, Alberto Calderone, Luisa Castagnoli, and Gianni Cesareni. Signor 2.0, the signaling network open resou… [cited by applicant]
Madalina Giurgiu, Julian Reinhard, Barbara Brauner, Irmtraud Dunger-Kaltenbach, Gisela Fobo, Goar Frishman, Corinna Montrone, and Andreas Ruepp. Corum: the comprehensive resource of mammalian protein complexes—2019. Nuc… [cited by applicant]
Marc Gillespie, Bijay Jassal, Ralf Stephan, Marija Milacic, Karen Rothfels, Andrea Senff-Ribeiro, Johannes Griss, Cristoffer Sevilla, Lisa Matthews, Chuqiao Gong, et al. The reactome pathway knowledgebase 2022. Nucleic … [cited by applicant]
Maria L Rizzo and Gábor J Székely. Energy distance. wiley interdisciplinary reviews: Computational statistics, 8(1):27-38, 2016. [cited by applicant]
Mark-Anthony Bray, Shantanu Singh, Han Han, Chadwick T Davis, Blake Borgeson, Cathy Hartland, Maria Kost-Alimova, Sigrun M Gustafsdottir, Christopher C Gibson, and Anne E Carpenter. Cell painting, a high-content image-b… [cited by applicant]
Michael F Cuccarese, Berton A Earnshaw, Katie Heiser, Ben Fogelson, Chadwick T Davis, Peter F McLean, Hannah B Gordon, Kathleen-Rose Skelly, Fiona L Weathersby, Vlad Rodic, Ian K Quigley, Elissa D Pastuzyn, Brandon M Me… [cited by applicant]
Nathan Lazar, et al. High-Resolution Genome-wide Mapping of Chromosome-arm-scale Truncations Induced by CRISPR-Cas9 Editing published in bioRxiv on Apr. 15, 2023. [cited by applicant]
Romain Lopez, Jeffrey Regier, Michael B Cole, Michael I Jordan, and Nir Yosef. Deep generative modeling for single-cell transcriptomics. Nature methods, 15(12):1053-1058, 2018. [cited by applicant]
U.S. Appl. No. 18/393,041, Feb. 23, 2024, Office Action. [cited by applicant]
Sivanandan et al. “A Pooled Cell Painting CRISPR Screening Platform Enables de novo Inference of Gene Function by Self-supervised Deep Learning”, https://www.biorxiv.org/contenU10.1101/2023.08.13.553051v3.abstract, bioR… [cited by applicant]
U.S. Appl. No. 18/392,989, Mar, 1, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/392,989, Apr. 29, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/393,041, May 3, 2024, Notice of Allowance. [cited by applicant]
Cited By (4)
US 12,651,432 US 12,657,939 US 12,718,527 US 12,738,081