IP Library Granted Patent US 12,373,949
Granted Patent B1
US 12,373,949 · App. 18/421,940 · Granted Jul 29, 2025

Auto-normalization for machine learning

Inventors: Ali Behrooz (South San Francisco, CA); Cheng-Hsun Wu (San Bruno, CA)
Assignee: Verily Life Sciences LLC
G06T7/0012G06N3/08G16H30/40G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,949
App. No.
18/421,940
Granted
Jul 29, 2025
Kind
B1
Abstract

Systems and methods for using a prediction model jointly with a normalization model to provide prediction results are provided. One example method includes receiving an input image of a tissue sample of a patient and generating a normalized image by applying a normalization model on the input image. The normalization model is configured to generate normalized data using input data for a prediction model, and the prediction model is configured to generate prediction results using normalized data generated by the normalization model. The normalization model and the prediction model are jointly trained. The method further includes generating a prediction of disease severity for the patient by applying the prediction model on the normalized image.

Claims (31)

1. A computer-implemented method, comprising:

receiving an input image of a tissue sample of a patient;

generating a normalized image by applying a normalization model on the input image, wherein the normalization model is configured to generate normalized data using input data for a prediction model, the prediction model configured to generate prediction results using normalized data generated by the normalization model, and wherein the normalization model and the prediction model are jointly trained by comparing a first set of prediction results generated by the prediction model using a first set of normalized training inputs generated by applying the normalization model to training inputs in a set of training samples for a first number of times and a second set of prediction results generated by the prediction model using a second set of normalized training inputs generated by applying the normalization model to the training inputs for a second number of times, the second number of times being different from the first number of times; and

generating a prediction for the patient by applying the prediction model on the normalized image.

2. The method of claim 1 , wherein the normalization model and the prediction model are jointly trained based on a loss function comprising an idempotence loss term measuring a difference between the first set of prediction results and the second set of prediction results.

3. The method of claim 2 , wherein the set of training samples comprise unlabeled samples that comprise the training inputs without corresponding training outputs, and wherein the idempotence loss term is calculated based on the training inputs in the unlabeled samples.

4. The method of claim 3 , wherein the set of training samples further comprises labeled samples that comprise training inputs and corresponding training outputs, and wherein the loss function further comprises a loss term for the labeled samples measuring a difference between prediction results generated by the prediction model and the normalization model based on the training inputs and the corresponding training outputs.

5. The method of claim 2 , wherein the idempotence loss term comprises a Kullback-Leibler divergence between the first set of prediction results and the second set of prediction results, a mean absolute error between the first set of prediction results and the second set of prediction results, or a mean square error between the first set of prediction results and the second set of prediction results.

6. The method of claim 2 , wherein the loss function further comprises an entropy loss term representing an entropy of the first set of prediction results or the second set of prediction results.

7. The method of claim 6 , wherein the input image is a pathology image and generating the prediction for the patient comprises detecting presence or absence of metastatic lesions in the pathology image using the prediction model.

8. A system comprising:

at least one processor; and

at least one non-transitory computer-readable medium comprising processor-executable instructions stored thereupon, which, when executed by the at least one processor, cause the processor to:

receive an input image of a tissue sample of a patient;

generate a normalized image by applying a normalization model on the input image, wherein the normalization model is configured to generate normalized data using input data for a prediction model, the prediction model configured to generate prediction results using normalized data generated by the normalization model, and wherein the normalization model and the prediction model are jointly trained by comparing a first set of prediction results generated by the prediction model using a first set of normalized training inputs generated by applying the normalization model to training inputs in a set of training samples for a first number of times and a second set of prediction results generated by the prediction model using a second set of normalized training inputs generated by applying the normalization model to the training inputs for a second number of times, the second number of times being different from the first number of times; and

generate a prediction for the patient by applying the prediction model on the normalized image.

9. The system of claim 8 , wherein the normalization model and the prediction model are jointly trained based on a loss function comprising an idempotence loss term measuring a difference between the first set of prediction results and the second set of prediction results.

10. The system of claim 9 , wherein the set of training samples comprise unlabeled samples that comprise the training inputs without corresponding training outputs, and wherein the idempotence loss term is calculated based on the training inputs in the unlabeled samples.

11. The system of claim 10 , wherein the set of training samples further comprises labeled samples that comprise training inputs and corresponding training outputs, and wherein the loss function further comprises a loss term for the labeled samples measuring a difference between prediction results generated by the prediction model and the normalization model based on the training inputs and the corresponding training outputs.

12. The system of claim 9 , wherein the idempotence loss term comprises a Kullback-Leibler divergence between the first set of prediction results and the second set of prediction results, a mean absolute error between the first set of prediction results and the second set of prediction results, or a mean square error between the first set of prediction results and the second set of prediction results.

13. The system of claim 9 , wherein the loss function further comprises an entropy loss term representing an entropy of the first set of prediction results or the second set of prediction results.

14. The system of claim 13 , wherein the input image is a pathology image and generating the prediction for the patient comprises detecting presence or absence of metastatic lesions in the pathology image using the prediction model.

15. A non-transitory computer-readable medium comprising processor-executable instructions to cause a processor to:

receive an input image of a tissue sample of a patient;

generate a normalized image by applying a normalization model on the input image, wherein the normalization model is configured to generate normalized data using input data for a prediction model, the prediction model configured to generate prediction results using normalized data generated by the normalization model, and wherein the normalization model and the prediction model are jointly trained by comparing a first set of prediction results generated by the prediction model using a first set of normalized training inputs generated by applying the normalization model to training inputs in a set of training samples for a first number of times and a second set of prediction results generated by the prediction model using a second set of normalized training inputs generated by applying the normalization model to the training inputs for a second number of times, the second number of times being different from the first number of times; and

generate a prediction for the patient by applying the prediction model on the normalized image.

16. The non-transitory computer-readable medium of claim 15 , wherein the normalization model and the prediction model are jointly trained based on a loss function comprising an idempotence loss term measuring a difference between the first set of prediction results and the second set of prediction results.

17. The non-transitory computer-readable medium of claim 16 , wherein the set of training samples comprise unlabeled samples that comprise the training inputs without corresponding training outputs, and wherein the idempotence loss term is calculated based on the training inputs in the unlabeled samples.

18. The non-transitory computer-readable medium of claim 17 , wherein the set of training samples further comprises labeled samples that comprise training inputs and corresponding training outputs, and wherein the loss function further comprises a loss term for the labeled samples measuring a difference between prediction results generated by the prediction model and the normalization model based on the training inputs and the corresponding training outputs.

19. The non-transitory computer-readable medium of claim 16 , wherein the idempotence loss term comprises a Kullback-Leibler divergence between the first set of prediction results and the second set of prediction results, a mean absolute error between the first set of prediction results and the second set of prediction results, or a mean square error between the first set of prediction results and the second set of prediction results.

20. The non-transitory computer-readable medium of claim 16 , wherein the loss function further comprises an entropy loss term representing an entropy of the first set of prediction results or the second set of prediction results.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2025
From: BEHROOZ, ALI; WU, CHENG-HSUN
To: VERILY LIFE SCIENCES LLC
Reel/Frame 070811/0619 →
CHANGE OF ADDRESS Recorded Nov 19, 2024
From: VERILY LIFE SCIENCES LLC
To: VERILY LIFE SCIENCES LLC
Reel/Frame 069390/0656 →
Continuity (2)
Continuation 17358621 · Jun 25, 2021
Provisional Application 62705403 · Jun 25, 2020
References Cited (28)
US 11042789B2 · Lee · 2021 [cited by examiner]
US 11915419B1 · Behrooz · 2024 [cited by examiner]
US 20160217368A1 · Ioffe et al. · 2016 [cited by applicant]
US 20180137642A1 · Malisiewicz et al. · 2018 [cited by applicant]
US 20180232883A1 · Sethi et al. · 2018 [cited by applicant]
US 20200311911A1 · Poole · 2020 [cited by examiner]
US 20220246301A1 · Choi · 2022 [cited by applicant]
CN 106919903B · 2019 [cited by applicant]
WO 2019081545A1 · 2019 [cited by applicant]
U.S. Appl. No. 17/358,621, Notice of Allowance, Oct. 23, 2023, 10 pages. [cited by applicant]
Badrinarayanan et al., “SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, No. 12, Dec. 2017, pp. 2481-2495. [cited by applicant]
Bejnordi et al., “Diagnostic Assessment of Deep Learning Algorithms for Detection of Lymph Node Metastases in Women With Breast Cancer”, Jama, vol. 318, No. 22, Dec. 12, 2017, pp. 2199-2210. [cited by applicant]
Chaurasia et al., “LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation”, IEEE Visual Communications and Image Processing (VCIP). IEEE, Jun. 14, 2017, 5 pages. [cited by applicant]
Chen et al., “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 4, May 12, 201… [cited by applicant]
Cubuk et al., “AutoAugment: Learning Augmentation Policies From Data”, Available Online at: arXiv preprint arXiv:1805.09501, Apr. 11, 2019, pp. 1-14. [cited by applicant]
Devries et al., “Improved Regularization of Convolutional Neural Networks with Cutout”, Available Online at: https://arxiv.org/pdf/1708.04552.pdf, Nov. 29, 2017, pp. 1-8. [cited by applicant]
Lin et al., “RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available Online at: arXiv:1611.06612, 20… [cited by applicant]
Paszke et al., “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation”, Available Online at: arXiv preprint arXiv:1606.02147, Jun. 7, 2016, pp. 1-10. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, International Conference on Medical Image Computing and Computer-Assisted Intervention, Available Online at URL: https://arxiv.org/p… [cited by applicant]
Salimans et al., “Improved Techniques for Training GANs”, Advances in Neural Information Processing Systems, vol. 29, Available Online at: https://arxiv.org/pdf/1606.03498.pdf, Jun. 10, 2016, pp. 1-10. [cited by applicant]
Valenza, “Linear Algebra: An Introduction to Abstract Mathematics”, Springer Science & Business Media, 2012, p. 22. [cited by applicant]
Wan et al., “Regularization of Neural Networks using DropConnect”, International Conference on Machine Learning, PMLR, 2013, 9 pages. [cited by applicant]
Wang et al., “EnAET: Self-Trained Ensemble Autoencoding Transformations for Semi-Supervised Learning”, Available Online at: arXiv preprint arXiv:1911.09265, Nov. 21, 2019, pp. 1-10. [cited by applicant]
Xie et al., “Unsupervised Data Augmentation for Consistency Training”, Available Online at: https://arxiv.org/pdf/1904.12848v4.pdf, 2019, pp. 1-19. [cited by applicant]
Zagoruyko et al., “Wide Residual Networks”, British Machine Vision Conference, 2016, pp. 1-12. [cited by applicant]
Zhang et al., “Mixup: Beyond Empirical Risk Minimization”, International Conference on Learning Representations 2018, Apr. 27, 2018, 13 pages. [cited by applicant]
Zhao et al., “Pyramid Scene Parsing Network”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available Online at: arXiv:1612.01105, 2017, pp. 2881-2890. [cited by applicant]
Zhong et al., “Random Erasing Data Augmentation”, Available Online at: arXiv preprint arXiv:1708.04896, Nov. 16, 2017, pp. 1-10. [cited by applicant]