IP Library › Granted Patent US 11,682,036
Granted Patent B2
US 11,682,036 · App. 17/197,511 · Granted Jun 20, 2023

Machine learning with data synthesization

Inventors: Robert Bryant Kaspar (Berkeley, CA); Alok Gupta (San Francisco, CA); Aman Dhesi (San Francisco, CA)
Assignee: DOORDASH, INC.
G06Q30/0244G06F18/217G06F18/2148G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,682,036
App. No.
17/197,511
Filed
Mar 10, 2021
Granted
Jun 20, 2023
Kind
B2
Art Unit
3682
USPC
705/14.43
Abstract

In some examples, a computing device may receive data from a plurality of groups of data sources. The computing device may create a training data set from a first portion of the received data and may create a plurality of validation data sets from a second portion of the received data. For example, each validation data set may correspond to a respective one of the groups of data sources. The computing device may train, using the training data set, a plurality of machine learning models configured for synthesizing data. For instance, respective ones of the machine learning models may correspond to respective ones of the groups of data sources. Further, the computing device may validate the respective machine learning models using the respective validation data set corresponding to the respective group to which the respective machine learning model being validated corresponds.

Claims (57)

1. A system comprising:

one or more processors configured by executable instructions to perform operations including:

receiving, by the one or more processors, data from a plurality of groups of data sources;

creating, by the one or more processors, a training data set from a first portion of the received data;

creating, by the one or more processors, a plurality of validation data sets from a second portion of the received data, each validation data set corresponding to a respective one of the groups of data sources and containing data received from that respective group exclusive of data received from other groups of the data sources;

training, by the one or more processors, using the training data set, a plurality of data synthetization machine learning models configured for synthesizing data, each respective one of the data synthetization machine learning models corresponding to a respective one of the groups of data sources;

validating, by the one or more processors, the respective data synthetization machine learning models using the respective validation data set corresponding to the respective group to which the respective data synthetization machine learning model being validated corresponds;

generating first synthetic data by inputting, to a first one of the data synthetization machine learning models, first data received from a first data source corresponding to a first group of data sources for which the first data synthetization machine learning model is trained; and

determining an allocation of resources based at least in part on the first data and the first synthetic data.

2. The system as recited in claim 1 , the validating comprising determining an optimal amount of synthetic data to produce, respectively, for individual data sources of the plurality of groups data sources.

3. The system as recited in claim 1 , the operations further comprising:

constructing a first curve using the first data received from the first data source and the first synthetic data generated by the first data synthetization machine learning model; and

determining the allocation of resources based at least in part on the first curve.

4. The system as recited in claim 3 , the operations further comprising:

generating second synthetic data by inputting, to a second one of the data synthetization machine learning models, second data received from a second data source corresponding to a second group of data sources for which the second data synthetization machine learning model is trained;

constructing, for the second group of the plurality of groups of data sources, a second curve using the second data received from the second data source and the second synthetic data generated by the second data synthetization machine learning model; and

determining the allocation of resources based at least in part on comparing respective slopes of the first curve and the second curve.

5. The system as recited in claim 1 , wherein the second portion of data is data received most recently within a past threshold period of time.

6. The system as recited in claim 1 , wherein the received data includes customer conversions, the operations further comprising determining the respective group to which to attribute a first customer conversion based on a touchpoint of a customer corresponding to the first customer conversion that occurred most recently prior to the first customer conversion.

7. A method comprising:

receiving, by one or more processors, data from a plurality of groups of data sources;

creating, by the one or more processors, a training data set from a first portion of the received data;

creating, by the one or more processors, a plurality of validation data sets from a second portion of the received data, each validation data set corresponding to a respective one of the groups of data sources;

training, by the one or more processors, using the training data set, a plurality of data synthetization machine learning models configured for synthesizing data, respective ones of the data synthetization machine learning models corresponding to respective ones of the groups of data sources;

validating, by the one or more processors, the respective data synthetization machine learning models using the respective validation data set corresponding to the respective group to which the respective data synthetization machine learning model being validated corresponds;

generating first synthetic data by inputting, to a first one of the data synthetization machine learning models, first data received from a first data source corresponding to a first group of data sources for which the first data synthetization machine learning model is trained; and

determining an allocation of resources based at least in part on the first data and the first synthetic data.

8. The method as recited in claim 7 , each validation data set corresponding to a respective one of the groups of data sources and containing data received from that respective group exclusive of data received from other groups of the data sources.

9. The method as recited in claim 7 , the validating comprising determining an optimal amount of synthetic data to produce, respectively, for individual data sources of the plurality of groups data sources.

10. The method as recited in claim 7 , further comprising:

constructing a first curve using the first data received from the first data source and the first synthetic data generated by the first data synthetization machine learning model; and

determining the allocation of resources based at least in part on the first curve.

11. The method as recited in claim 10 , further comprising:

generating second synthetic data by inputting, to a second one of the data synthetization machine learning models, second data received from a second data source corresponding to a second group of data sources for which the second data synthetization machine learning model is trained;

constructing, for the second group of the plurality of groups of data sources, a second curve using the second data received from the second data source and the second synthetic data generated by the second data synthetization machine learning model; and

determining the allocation of resources based at least in part on comparing respective slopes of the first curve and the second curve.

12. The method as recited in claim 11 , further comprising determining a plurality of bids to send to a plurality of service provider computing devices based at least in part on the allocation of resources.

13. The method as recited in claim 7 , wherein the received data includes customer conversions, the method further comprising determining the respective group to which to attribute a first customer conversion based on a touchpoint of a customer corresponding to the first customer conversion that occurred most recently prior to the first customer conversion.

14. A non-transitory computer-readable medium maintaining instructions executable to configure one or more processors to perform operations comprising:

receiving data from a plurality of groups of data sources;

creating a training data set from a first portion of the received data;

creating a plurality of validation data sets from a second portion of the received data, each validation data set corresponding to a respective one of the groups of data sources;

training, using the training data set, a plurality of data synthetization machine learning models configured for synthesizing data, respective ones of the data synthetization machine learning models corresponding to respective ones of the groups of data sources;

validating the respective data synthetization machine learning models using the respective validation data set corresponding to the respective group to which the respective data synthetization machine learning model being validated corresponds;

generating first synthetic data by inputting, to a first one of the data synthetization machine learning models, first data received from a first data source corresponding to a first group of data sources for which the first data synthetization machine learning model is trained; and

determining an allocation of resources based at least in part on the first data and the first synthetic data.

15. The non-transitory computer-readable medium as recited in claim 14 , each validation data set corresponding to a respective one of the groups of data sources and containing data received from that respective group exclusive of data received from other groups of the data sources.

16. The non-transitory computer-readable medium as recited in claim 14 , the validating comprising determining an optimal amount of synthetic data to produce, respectively, for individual data sources of the plurality of groups data sources.

17. The non-transitory computer-readable medium as recited in claim 14 , the operations further comprising:

constructing a first curve using the first data received from the first data source and the first synthetic data generated by the first data synthetization machine learning model; and

determining the allocation of resources based at least in part on the first curve.

18. The non-transitory computer-readable medium as recited in claim 17 , the operations further comprising:

generating second synthetic data by inputting, to a second one of the data synthetization machine learning models, second data received from a second data source corresponding to a second group of data sources for which the second data synthetization machine learning model is trained;

constructing, for the second group of the plurality of groups of data sources, a second curve using the second data received from the second data source and the second synthetic data generated by the second data synthetization machine learning model; and

determining the allocation of resources based at least in part on comparing respective slopes of the first curve and the second curve.

19. The non-transitory computer-readable medium as recited in claim 18 , the operations further comprising determining a plurality of bids to send to a plurality of service provider computing devices based at least in part on the allocation of resources.

20. The non-transitory computer-readable medium as recited in claim 14 , wherein the received data includes customer conversions, the operations further comprising determining the respective group to which to attribute a first customer conversion based on a touchpoint of a customer corresponding to the first customer conversion that occurred most recently prior to the first customer conversion.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: KASPAR, ROBERT BRYANT; GUPTA, ALOK; DHESI, AMAN
To: DOORDASH, INC.
Reel/Frame 055549/0804 →
Continuity (1)
Related Publication 20220292542A1 · Sep 15, 2022