IP Library › Granted Patent US 12,361,731
Granted Patent B2
US 12,361,731 · App. 17/911,841 · Granted Jul 15, 2025

Systems and methods for predicting expression levels

Inventors: Anne-Laure Bauchet (Paris, FR); Anthony Mei (Bridgewater, NJ); Qi Tang (Bridgewater, NJ)
Assignee: Sanofi
G06V20/695G06T7/0012G06T7/11G06V10/25G06V10/75G06V10/82G06V20/698G16H50/20G06T2207/20021G06T2207/20081G06T2207/20084G06T2207/30096G06T2207/30204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,731
App. No.
17/911,841
Granted
Jul 15, 2025
Kind
B2
Abstract

One or more methods of predicting expression levels. At least one of the methods includes preprocessing image data representing at least one biological image of a patient to generate preprocessed image data representing at least one preprocessed biological image of the patient; and applying a trained machine learning model to the preprocessed image data to predict, based at least partially on the at least one preprocessed biological image, an expression level of a biological indicator.

Claims (58)

1. A method performed by one or more computers, the method comprising:

obtaining image data representing at least one hematoxylin and eosin (H&E)-stained biological image of a patient; and

processing a model input that comprises the image data representing the at least one H&E-stained biological image using a trained machine learning model to generate, as a model output of the machine learning model, a predicted aggregate expression level of a protein across tissue that is depicted by the at least one H&E-stained biological image, wherein the protein is a biomarker for a particular disease;

wherein the machine learning model has been trained by performing operations comprising:

obtaining a set of training examples that each include: (i) at least one training H&E-stained biological image; and (ii) a corresponding target aggregate protein expression level across tissue that is depicted by the at least one training H&E-stained biological image; and

training the machine learning model on the set of training examples, comprising, for each training example, training the machine learning model to reduce a discrepancy between: (i) a predicted aggregate protein expression level generated by processing the at least one training H&E-stained biological image of the training example using the machine learning model, and (ii) the target aggregate protein expression level specified by the training example; and

outputting data identifying the predicted aggregate expression level of the protein.

2. The method of claim 1 , further comprising:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level exceeds the threshold expression level, to perform an immunohistochemistry (IHC) screening test for the patient.

3. The method of claim 2 , further comprising:

performing the IHC screening test for the patient to determine an IHC score; and

determining, based at least partially on the determined IHC score, to enroll the patient in a clinical trial.

4. The method of claim 1 , further comprising:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level does not exceed the threshold expression level, to not perform an IHC screening test for the patient.

5. The method of claim 1 , further comprising preprocessing the image data by performing an automatic image thresholding process to segment one or more tissue regions of the at least one H&E-stained biological image.

6. The method of claim 1 , further comprising preprocessing the image data by separating the at least one H&E-stained biological image into a plurality of image tiles.

7. The method of claim 6 , wherein preprocessing the image data comprises separating each of the plurality of image tiles into a plurality of sub-tiles.

8. A system, comprising:

a computer-readable memory comprising computer-executable instructions; and

at least one processor configured to execute the computer-executable instructions, wherein when the at least one processor is executing the computer-executable instructions, the at least one processor is configured to carry out operations comprising:

obtaining image data representing at least one hematoxylin and eosin (H&E)-stained biological image of a patient; and

processing a model input that comprises the image data representing the at least one H&E-stained biological image using a trained machine learning model to generate, as a model output of the machine learning model, a predicted aggregate expression level of a protein across tissue that is depicted by the at least one H&E-stained biological image, wherein the protein is a biomarker for a particular disease;

wherein the machine learning model has been trained by performing operations comprising:

obtaining a set of training examples that each include: (i) at least one training H&E-stained biological image; and (ii) a corresponding target aggregate protein expression level across tissue that is depicted by the at least one training H&E-stained biological image; and

training the machine learning model on the set of training examples, comprising, for each training example, training the machine learning model to reduce a discrepancy between: (i) a predicted aggregate protein expression level generated by processing the at least one training H&E-stained biological image of the training example using the machine learning model, and (ii) the target aggregate protein expression level specified by the training example; and

outputting data identifying the predicted aggregate expression level of the protein.

9. The system of claim 8 , the operations further comprising:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level exceeds the threshold expression level, to recommend performance of an immunohistochemistry (IHC) screening test for the patient.

10. The system of claim 9 , the operations further comprising:

performing the IHC screening test for the patient to determine an IHC score; and

determining, based at least partially on the determined IHC score, to enroll the patient in a clinical trial.

11. The system of claim 8 , the operations further comprising:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level does not exceed the threshold expression level, to recommend against performance of an IHC screening test for the patient.

12. The system of claim 8 , wherein the operations further comprise preprocessing the image data by performing an automatic image thresholding process to segment one or more tissue regions of the at least one H&E-stained biological image.

13. The system of claim 8 , further comprising preprocessing the image data by separating the at least one H&E-stained biological image into a plurality of image tiles.

14. The system of claim 13 , wherein preprocessing the image data comprises separating each of the plurality of image tiles into a plurality of sub-tiles.

15. One or more non-transitory computer-readable storage media storing instructions executable by one or more processors to cause the processors to perform operations comprising:

obtaining image data representing at least one hematoxylin and eosin (H&E)-stained biological image of a patient; and

processing a model input that comprises the image data representing the at least one H&E-stained biological image using a trained machine learning model to generate, as a model output of the machine learning model, a predicted aggregate expression level of a protein across tissue that is depicted by the at least one H&E-stained biological image, wherein the protein is a biomarker for a particular disease;

wherein the machine learning model has been trained by performing operations comprising:

obtaining a set of training examples that each include: (i) at least one training H&E-stained biological image; and (ii) a corresponding target aggregate protein expression level across tissue that is depicted by the at least one training H&E-stained biological image; and

training the machine learning model on the set of training examples, comprising, for each training example, training the machine learning model to reduce a discrepancy between: (i) a predicted aggregate protein expression level generated by processing the at least one training H&E-stained biological image of the training example using the machine learning model, and (ii) the target aggregate protein expression level specified by the training example; and

outputting data identifying the predicted aggregate expression level of the protein.

16. The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level exceeds the threshold expression level, to perform an immunohistochemistry (IHC) screening test for the patient.

17. The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise:

determining whether the predicted aggregate expression level of the protein exceeds a threshold expression level; and

determining, in response to determining that the predicted aggregate expression level does not exceed the threshold expression level, to not perform an IHC screening test for the patient.

18. The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise:

preprocessing the image data by performing an automatic image thresholding process to segment one or more tissue regions of the at least one H&E-stained biological image.

19. The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise:

preprocessing the image data by separating the at least one H&E-stained biological image into a plurality of image tiles.

20. The one or more non-transitory computer-readable storage media of claim 15 , wherein the machine learning model comprises a deep neural network.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2023
From: MEI, ANTHONY; TANG, QI
To: SANOFI
Reel/Frame 062443/0050 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2023
From: BAUCHET, ANNE-LAURE
To: SANOFI-AVENTIS RECHERCHE & DEVELOPMENT
Reel/Frame 062443/0119 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2023
From: SANOFI-AVENTIS RECHERCHE & DEVELOPMENT
To: SANOFI
Reel/Frame 062443/0162 →
Priority Claims (1)
EP 20315297 · Jun 5, 2020 · regional
Continuity (2)
Provisional Application 62991412 · Mar 18, 2020
Related Publication 20230143701A1 · May 11, 2023
References Cited (38)
US 11783604B2 · Peng · 2023 [cited by examiner]
US 20180330500A1 · Sakurai et al. · 2018 [cited by applicant]
US 20190150857A1 · Nye et al. · 2019 [cited by applicant]
US 20190206056A1 · Georgescu · 2019 [cited by examiner]
US 20190347557A1 · Khan · 2019 [cited by examiner]
US 20190355113A1 · Wirch · 2019 [cited by examiner]
US 20200321102A1 · Barnes et al. · 2020 [cited by applicant]
US 20200394825A1 · Stumpe · 2020 [cited by examiner]
US 20210073986A1 · Kapur · 2021 [cited by examiner]
US 20210279875A1 · Saur · 2021 [cited by examiner]
US 20210312620A1 · Zuo · 2021 [cited by examiner]
US 20220130041A1 · Dogdas · 2022 [cited by examiner]
US 20230419694A1 · Stumpe · 2023 [cited by examiner]
AU 2020421604A1 · 2022 [cited by examiner]
CN 111542830A · 2020 [cited by examiner]
EP 3729369A2 · 2020 [cited by examiner]
EP 3639191B1 · 2021 [cited by examiner]
JP 2018187384A · 2018 [cited by applicant]
JP 2019093137A · 2019 [cited by applicant]
WO WO2014102130 · 2014 [cited by applicant]
WO WO2014102130A1 · 2014 [cited by examiner]
WO WO2018115055 · 2018 [cited by applicant]
WO WO2018115055A1 · 2018 [cited by examiner]
WO WO2019121564 · 2019 [cited by applicant]
WO WO2019121564A2 · 2019 [cited by examiner]
WO WO2019172901A1 · 2019 [cited by examiner]
WO WO2019235828A1 · 2019 [cited by examiner]
WO WO2020014477A1 · 2020 [cited by examiner]
WO WO2021133847A1 · 2021 [cited by examiner]
Bychkov et al., “Deep learning based tissue analysis predicts outcome in colorectal cancer,” Scientific Reports, Feb. 21, 2018, 8:3395, 11 pages. [cited by applicant]
DiMasi et al., “Innovation in the pharmaceutical industry: New estimates of R&D costs,” Journal of Health Economics, May 2016, 47:20-33. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2021/022675, mailed Sep. 29, 2022, 11 page. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2021/022675, mailed on Aug. 12, 2021, 16 pages. [cited by applicant]
Otsu, “A Threshold Selection Method from Gray-Level Histograms,” IEEE Transactions on Systems, Man, and Cybernetics, Jan. 1979, SMC-9(1):62-66. [cited by applicant]
Sha et al., “Multi-Field-of-View Deep Learning Model Predicts Nonsmall Cell Lung Cancer Programmed Death-Ligand 1 Status from Whole-Slide Hematoxylin and Eosin Images,” Journal of Pathology Informatics, Jul. 23, 2019, 1… [cited by applicant]
Shin et al., “Addressing the challenges of applying precision oncology,” NPJ Precision Oncology, Sep. 4, 2017, 28:1-10. [cited by applicant]
Vamathevan et al., “Applications of machine learning in drug discovery and development,” Nature Reviews Drug Discovery, Apr. 11, 2019, 18(6):463-477. [cited by applicant]
Xu et al., “GAN-based virtual re-staining: a promising solution for whole slide image analysis,” CoRR, Submitted on Jan. 13, 2019, arXiv: 1901.04059v1, 16 pages. [cited by applicant]