IP Library › Granted Patent US 12,450,735
Granted Patent B2
US 12,450,735 · App. 17/901,397 · Granted Oct 21, 2025

Learning representations of nuclei in histopathology images with contrastive loss

Inventors: Chao Feng (New York, NY); Chad Vanderbilt (New York, NY); Thomas Fuchs (New York, NY)
Assignee: Memorial Sloan-Kettering Cancer Center
G06T7/0012G06T3/60G06T5/70G06V10/25G06V10/762G06V10/771
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,735
App. No.
17/901,397
Granted
Oct 21, 2025
Kind
B2
Abstract

Presented herein are systems and methods for classifying features from biomedical images. A computing system may identify a first portion corresponding to an ROI in a first biomedical image derived from a sample. The ROI of the first biomedical image may correspond to a feature of the sample. The computing system may generate a first embedding vector using the first portion of the first biomedical image. The computing system may apply the first embedding vector to a clustering model. The clustering model may have a feature space to define a plurality of conditions. The clustering model may be trained using a second embedding vectors generated from a corresponding second portions with at least one of a plurality of image transformation. The computing system may determine a condition for the feature based on applying the first embedding vector to the clustering model.

Claims (43)

1. A method of classifying regions of interest (ROIs) in biomedical images, comprising:

identifying, by a computing system, a first portion corresponding to an ROI in a first biomedical image derived from a sample, the ROI of the first biomedical image corresponding to a feature of the sample;

generating, by the computing system, a first embedding vector using the first portion of the first biomedical image;

providing, by the computing system, the first embedding vector to a clustering model, the clustering model having a feature space to define a plurality of conditions,

wherein the clustering model is trained using a second plurality of embedding vectors generated from a corresponding second plurality of portions, each portion of the second plurality of portions corresponding to a second ROI in a second biomedical image with at least one of a plurality of image transformations, and updating the feature space based on a plurality of positions corresponding to the plurality of second embedding vectors in accordance with at least one of (i) a first loss metric to align a first subset of embedding vectors generated from a corresponding first subset of a plurality of second portions and (ii) a second loss metric to disperse the first subset of embedding vectors from at least one second subset of embedding vectors generated from a corresponding second subset of second portions;

determining, by the computing system, a condition for the feature from the plurality of conditions based on applying the first embedding vector to the clustering model; and

storing, by the computing system, in one or more data structures, an association between the condition for the feature and the first biomedical image.

2. The method of claim 1 , further comprising providing, by the computing system, information based on the association between the condition for the first feature and the first biomedical image.

3. The method of claim 1 , wherein determining the condition further comprises identifying, from a plurality of regions defined in the feature space for the plurality of conditions, a region in which the embedding vector is situated.

4. The method of claim 1 , wherein the plurality of conditions defined by the feature space further comprises (i) a first subset of conditions as during training and (ii) a second subset of conditions subsequent to the training.

5. The method of claim 1 , wherein the plurality of image transformations comprises at least one of color jittering, blurring, rotation, flipping, or background replacement.

6. The method of claim 1 , wherein identifying the first portion further comprises generating a bounding box to define the ROI within the biomedical image.

7. The method of claim 1 , wherein the feature includes at least one nucleus in the sample and the plurality of conditions includes a plurality of cancer subtypes.

8. A method of training models to classify regions of interest (ROIs) in biomedical images, comprising:

identifying, by a computing system, a training dataset identifying a plurality of instances, each of the plurality of instances comprising a first portion corresponding to a respective ROI in a biomedical image derived from a sample, the respective ROI corresponding to a feature in the sample;

adding, by the computing system, at least one of a plurality of image transformations to the first portion to generate a plurality of second portions for each instance of the plurality of instances;

generating, by the computing system, a plurality of embedding vectors from the plurality of second portions for each instance of the plurality of instances;

providing, by the computing system, the plurality of embedding vectors to a clustering model to determine a plurality of positions within a feature space defined by the clustering model for a plurality of conditions;

determining, by the computing system, (i) a first loss metric to align a first subset of embedding vectors generated from a corresponding first subset of a plurality of second portions and (ii) a second loss metric to disperse the first subset of embedding vectors from at least one second subset of embedding vectors generated from a corresponding second subset of second portions;

updating, by the computing system, the feature space of the clustering model based on the plurality of positions corresponding to the plurality of embedding vectors in accordance with at least one of the first loss metric or the second loss metric; and

storing, by the computing system, the feature space of the clustering model to define the plurality of conditions.

9. The method of claim 8 ,

wherein updating the feature space further comprises updating the feature space in accordance with the first loss metric and the second loss metric.

10. The method of claim 8 , wherein updating the feature space further comprises:

determining a distance between an embedding vector of the plurality of embedding vectors and a first region of a plurality of regions within the feature space corresponding to the plurality of conditions defined by the training dataset; and

defining, responsive to the distance being greater than a threshold, a second region to include to the plurality of regions corresponding to a condition different from any of the plurality of conditions.

11. The method of claim 8 , wherein identifying the respective portion in each instance of the plurality of instances further comprises generating a bounding box to define the ROI within the biomedical image.

12. The method of claim 8 , wherein adding at least one of the plurality of image transformations further comprises selecting, for the first portion, an image transformation from the plurality of image transformations in accordance with a function.

13. The method of claim 8 , wherein the plurality of image comprises at least one of color jittering, blurring, rotation, flipping, or background replacement.

14. The method of claim 8 , wherein the feature includes at least one nucleus in the sample and the plurality of conditions include a plurality of cancer subtypes.

15. A system for classifying regions of interest (ROIs) in biomedical images, comprising:

a computing system having one or more processors coupled with memory, configured to:

identify a first portion corresponding to an ROI in a first biomedical image derived from a sample, the ROI of the first biomedical image corresponding to a feature of the sample;

generate a first embedding vector using the first portion of the first biomedical image;

provide the first embedding vector to a clustering model, the clustering model having a feature space to define a plurality of conditions,

wherein the clustering model is trained using a second plurality of embedding vectors generated from a corresponding second plurality of portions, each portion of the second plurality of portions corresponding to a second ROI in a second biomedical image with at least one of a plurality of image transformations, and updating the feature space based on a plurality of positions corresponding to the plurality of second embedding vectors in accordance with at least one of (i) a first loss metric to align a first subset of embedding vectors generated from a corresponding first subset of a plurality of second portions and (ii) a second loss metric to disperse the first subset of embedding vectors from at least one second subset of embedding vectors generated from a corresponding second subset of second portions;

determine a condition for the feature from the plurality of conditions based on applying the first embedding vector to the clustering model; and

store, in one or more data structures, an association between the condition for the feature and the first biomedical image.

16. The system of claim 15 , wherein the computing system is further configured to provide information based at least on the association between the condition for the first feature and the first biomedical image.

17. The system of claim 15 , wherein the computing system is further configured to determining the condition by identifying, from a plurality of regions defined in the feature space for the plurality of conditions, a region in which the embedding vector is situated.

18. The system of claim 15 , wherein the plurality of conditions defined by the feature space further comprises (i) a first subset of conditions during training and (ii) a second subset of conditions subsequent to the training.

19. The system of claim 15 , wherein the plurality of image transformations comprises at least one of color jittering, blurring, rotation, flipping, or background replacement.

20. The system of claim 15 , wherein the feature includes at least one nucleus in the sample and the plurality of conditions include a plurality of cancer subtypes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2023
From: FENG, CHAO; VANDERBILT, CHAD; FUCHS, THOMAS
To: MEMORIAL SLOAN-KETTERING CANCER CENTER
Reel/Frame 065797/0809 →
Continuity (2)
Provisional Application 63240177 · Sep 2, 2021
Related Publication 20230070874A1 · Mar 9, 2023
References Cited (29)
US 20100329529A1 · Feldman · 2010 [cited by examiner]
US 20220164946A1 · Pati · 2022 [cited by examiner]
US 20230170050A1 · Cooper · 2023 [cited by examiner]
US 20230419695A1 · Akers · 2023 [cited by examiner]
US 20240016901A1 · Sexton · 2024 [cited by examiner]
Yang et al., “Applying Deep Neural Network Analysis to High-Content Image-Based Assays,” SLAS Discovery 2019, vol. 24(8) 829-841 (Year: 2019). [cited by examiner]
Doyle et al., “Automated grading of breast cancer histopathology using spectral clustering with textural and architectural image features,” 2008 5th IEEE International Symposium on Biomedical Imaging: From Nano to Macro… [cited by examiner]
Basavanhally et al., “Computerized Image-Based Detection and Grading of Lymphocytic Infiltration in HER2+ Breast Cancer Histopathology,” IEEE Transactions on Biomedical Engineering, vol. 57, No. 3, Mar. 2010 (Year: 2010… [cited by examiner]
Abadi, Martín, et al., “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems”, Preliminary White Paper, 2015, 19 pages. [cited by applicant]
Levy-Jurgenson, Alona, et al., “Spatial transcriptomics inferred from pathology whole-slide images links tumor heterogeneity to survival in breast and lung cancer,” Scientific reports, 10(1), 2020, 11 pages. [cited by applicant]
Borros Arneth, “Tumor Microenvironment”, Medicina, 56(1):15, 2020, 21 pages. [cited by applicant]
Xie, Chensu, et al., “VOCA: Cell Nuclei Detection in Histopathology Images by Vector Oriented Confidence Accumulation”, Proceedings of Machine Learning Research, pp. 527-539, MIDL 2019, 13 pages. [cited by applicant]
Müllner, Daniel, “fastcluster: Fast Hierarchical, Agglomerative Clustering Routines for R and Python,” Journal of Statistical Software, 53(9), 2013, 18 pages, URL http://www.jstatsoft. org/v53/i09/. [cited by applicant]
Diao, James A., et al., “Dense, high-resolution mapping of cells and tissues from pathology images for the interpretable prediction of molecular phenotypes in cancer,” bioRxiv preprint, 2020, 14 pages. [cited by applicant]
Johnson, Jeff, et al., “Billion-scale similarity search with GPUs”, arXiv preprint: arXiv:1702.08734v1, 2017, 12 pages. [cited by applicant]
Gamper, Jevgenij, et al., “PanNuke Dataset Extension, Insights and Baselines”, arXiv:2003.10778v7, 2020, 12 pages. [cited by applicant]
Abduljabbar, Khalid, et al., “Geospatial immune variability illuminates differential evolution of lung adenocarcinoma”, Nature Medicine, Author manuscript, available in PMC May 2, 2021, 45 pages. [cited by applicant]
Sirinukunwattana, Korsuk, et al., “Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images”, IEEE Transactions on Medical Imaging, 35(5):1196-1206, 2016. [cited by applicant]
McInnes, Leland, et al., “hdbscan: Hierarchical density based clustering”, The Journal of Open Source Software, 2(11), 2017. doi: 10.21105/joss.00205. [cited by applicant]
Paszke, Adam, et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, URL http://papers.neurips.cc/pape… [cited by applicant]
Campello, Ricardo, J.G.B., et al., “Density-Based Clustering Based on Hierarchical Density Estimates”, In Pacific-Asia conference on knowledge discovery and data mining, PAKDD 2013. Lecture Notes in Computer Science, vo… [cited by applicant]
Campello, Ricardo, J.G.B., et al., “A framework for semi-supervised and unsupervised optimal extraction of clusters from hierarchies”, Data Mining and Knowledge Discovery, 27(3):344-371, 2013. [cited by applicant]
Graham, Simon, et al., “HoVer-Net: Simultaneous Segmentation and Classification of Nuclei in Multi-Tissue Histology Images”, Medical Image Analysis, 58:101563, 2019, arXiv:1812.06499v5. [cited by applicant]
Paijens, Sterre T., et al., “Tumor-infiltrating lymphocytes in the immunotherapy era”, Cellular & Molecular Immunology, 2020, 18 pages. [cited by applicant]
Chen, Ting, et al., “A Simple Framework for Contrastive Learning of Visual Representations”, Proceedings of the 37 [cited by applicant]
Wang, Tongzhou, et al., “Understanding Contrastive Representation Learning Through Alignment and Uniformity on the Hypersphere,” Proceedings of the 37 [cited by applicant]
Van Gansbeke, Wouter, et al., “SCAN: Learning to Classify Images without Labels”, arXiv preprint arXiv:2005.12320v2, 2020. [cited by applicant]
Chen, Xinlei, et al., “Improved Baselines with Momentum Contrastive Learning”, arXiv preprint arXiv:2003.04297v1, 2020. [cited by applicant]
Zhou, Yanning, et al., “CGC-Net: Cell Graph Convolutional Network for Grading of Colorectal Cancer Histology Images,” In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, arXiv:1909.0106… [cited by applicant]