IP Library › Granted Patent US 11,544,796
Granted Patent B1
US 11,544,796 · App. 16/599,607 · Granted Jan 3, 2023

Cross-domain machine learning for imbalanced domains

Inventors: Moustafa Abdalla Mohamed (Kirkland, WA); Lifan Chen (Mercer Island, WA); Sandesh Govind Shridhar (Bellevue, WA); Raghava Gupta Valiveti (Bangalore, IN)
Assignee: AMAZON TECHNOLOGIES, INC.
G06Q40/10G06F17/16G06K9/6215G06K9/6223G06N3/0454G06N20/00G06Q10/0831G06Q20/207G06Q30/0283
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,796
App. No.
16/599,607
Granted
Jan 3, 2023
Kind
B1
Abstract

Devices and techniques are generally described for cross-domain machine learning. A first machine learning model may be trained using first data of a first domain. Predictions may be generated by inputting a plurality of domain data from other domains apart from the first domain into the first machine learning model. For each of the predictions, a prediction error may be determined. A grouping of similar domains from among the other domains may be determined based on the prediction errors. A second machine learning model may be trained for the grouping of similar domains.

Claims (82)

1. A computer-implemented method of cross-domain machine learning, the method comprising:

sending first data of a first domain to a first machine learning model trained using second data of a second domain;

determining a first error describing a difference between predicted values generated by the first machine learning model and actual values of the first data of the first domain;

sending the second data of the second domain to a second machine learning model trained using the first data of the first domain;

determining a second error describing a difference between predicted values generated by the second machine learning model and actual values of the second data of the second domain;

determining a similarity between the first domain and the second domain based on the first error and the second error;

training a third machine learning model using the first data of the first domain and the second data of the second domain;

receiving third data of the first domain; and

generating a prediction for the third data using the third machine learning model.

2. The computer-implemented method of claim 1 , further comprising:

generating, using the first error and the second error, a similarity matrix describing similarities among the first domain and the second domain.

3. The computer-implemented method of claim 1 , further comprising:

determining a plurality of groupings of domains; and

generating, for each of the plurality of groupings of domains, a respective fourth machine learning model.

4. The computer-implemented method of claim 1 , further comprising:

determining a first number of samples of the first data of the first domain;

determining a total number of samples across a plurality of domains including the first domain and the second domain; and

determining a sample weight using a ratio of the first number of samples to the total number of samples, wherein the training the third machine learning model comprises applying the sample weight to a loss function of the third machine learning model.

5. The computer-implemented method of claim 1 , further comprising:

determining a first optimized hyper-parameter for the first machine learning model; and

determining a second optimized hyper-parameter for the second machine learning model.

6. The computer-implemented method of claim 5 , further comprising:

training the first machine learning model using the first data of the first domain and the second data of the second domain; and

training the second machine learning model using the first data of the first domain and the second data of the second domain.

7. The computer-implemented method of claim 1 , further comprising:

receiving fourth data;

determining that the fourth data corresponds to the second domain;

sending the fourth data to the third machine learning model; and

generating, by inputting the fourth data into the third machine learning model, a prediction for the fourth data.

8. A system comprising:

at least one processor; and

non-transitory computer-readable memory storing instructions that, when executed by the at least one processor, are effective to:

send first data of a first domain to a first machine learning model trained using second data of a second domain;

determine a first error describing a difference between predicted values generated by the first machine learning model and actual values of the first data of the first domain;

send the second data of the second domain to a second machine learning model trained using the first data of the first domain;

determine a second error describing a difference between predicted values generated by the second machine learning model and actual values of the second data of the second domain;

determine a similarity between the first domain and the second domain based on the first error and the second error;

train a third machine learning model using the first data of the first domain and the second data of the second domain;

receive third data of the first domain; and

generate a prediction for the third data using the third machine learning model.

9. The system of claim 8 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

generate, using the first error and the second error, a similarity matrix describing similarities among the first domain and the second domain.

10. The system of claim 8 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

determine a plurality of groupings of domains; and

generate, for each of the plurality of groupings of domains, a respective fourth machine learning model.

11. The system of claim 8 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

determine a first number of samples of the first data of the first domain;

determine a total number of samples across a plurality of domains including the first domain and the second domain; and

determine a sample weight using a ratio of the first number of samples to the total number of samples, wherein the training the third machine learning model comprises applying the sample weight to a loss function of the third machine learning model.

12. The system of claim 8 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

determine a first optimized hyper-parameter for the first machine learning model; and

determine a second optimized hyper-parameter for the second machine learning model.

13. The system of claim 12 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

train the first machine learning model using the first data of the first domain and the second data of the second domain; and

train the second machine learning model using the first data of the first domain and the second data of the second domain.

14. The system of claim 8 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

receive fourth data;

determine that the fourth data corresponds to the second domain;

send the fourth data to the third machine learning model; and

generate, by inputting the fourth data into the third machine learning model, a prediction for the fourth data.

15. A computer-implemented method comprising:

sending first data of a first domain to a first machine learning model trained using second data of a second domain;

determining a first error describing a difference between predicted values generated by the first machine learning model and actual values of the first data of the first domain;

determining a similarity between the first domain and the second domain based on the first error and a second error associated with a second machine learning model trained using the first data of the first domain;

training a third machine learning model using the first data of the first domain and the second data of the second domain;

receiving third data of the first domain; and

generating a prediction for the third data using the third machine learning model.

16. The computer-implemented method of claim 15 , further comprising:

generating, using the first error and the second error, a similarity matrix describing similarities among the first domain and the second domain.

17. The computer-implemented method of claim 15 , further comprising:

determining a plurality of groupings of domains; and

generating, for each of the plurality of groupings of domains, a respective fourth machine learning model.

18. The computer-implemented method of claim 15 , further comprising:

determining a first number of samples of the first data of the first domain;

determining a total number of samples across a plurality of domains including the first domain and the second domain; and

determining a sample weight using a ratio of the first number of samples to the total number of samples, wherein the training the third machine learning model comprises applying the sample weight to a loss function of the third machine learning model.

19. The computer-implemented method of claim 15 , further comprising:

determining a first optimized hyper-parameter for the first machine learning model; and

determining a second optimized hyper-parameter for the second machine learning model.

20. The computer-implemented method of claim 19 , further comprising:

training the first machine learning model using the first data of the first domain and the second data of the second domain; and

training the second machine learning model using the first data of the first domain and the second data of the second domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2019
From: MOHAMED, MOUSTAFA ABDALLA; CHEN, LIFAN; SANDESH GOVIND SHRIDHAR, .; VALIVETI, RAGHAVA GUPTA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050690/0446 →
Cited By (1)
US 12,688,523