Machine learning techniques for content delivery service
A system is disclosed to host a first machine learning model and a second machine learning model, where the first machine learning model is trained using a first privacy-compliant dataset, and the second machine learning model is trained using a second privacy-compliant dataset. The system is to obtain a digital content request that includes a dataset including user data and digital content data. The dataset is processed using the first machine learning model to generate a first score representing a probability of a user engaging with digital content and using the second machine learning model to generate a second score representing a similar probability. Based on one or both of these scores, the system selects at least one digital content from a plurality of digital content.
1 . A computer-implemented method, comprising:
providing a data clean room to host first data from an external entity and second data of an entity that provides the data clean room, the data clean room configured to anonymize the first and second data to generate first anonymized data and second anonymized data, the data clean room further configured to at least restrict reading the first data and the second data based on one or more predefined permissions associated with the data clean room;
providing a service to host at least a first trained machine learning (ML) model and a second trained ML model, the first ML model trained using user data, digital content data, and contextual data determined by the external entity and at least partially comprised in the first anonymized data and the second trained ML model trained using user data, digital content data, and contextual data determined by the entity and at least partially comprised in the second anonymized data;
obtaining a digital content request triggered based on a digital content opportunity on a computing device, the digital content request comprising an information dataset including information associated with the computing device, anonymized user information, and contextual information;
processing the information dataset using the first trained ML model to generate a first propensity score representing a first probability of a user engaging digital content;
processing the information dataset using the second trained ML model to generate a second propensity score representing a second probability of a user engaging digital content; and
based, at least in part, on one or more of the first and second propensity scores, selecting at least one digital content from a plurality of digital content to display on the computing device.
2 . The computer-implemented method of claim 1 , wherein selecting the at least one digital content is based on the first and second propensity scores.
3 . The computer-implemented method of claim 1 , wherein the external entity is a first company and the entity is a second company, the second company providing a demand side service to make a bid, on behalf of the first company, to display the at least one digital content on the computing device.
4 . The computer-implemented method of claim 1 , wherein the second trained ML model is based on an untrained ML model provided by the entity, the untrained ML model comprising adjustable input features comprising one or more of user demographics, digital content format, user browsing history, device type, or historical click through rate (CTR).
5 . A system, comprising:
one or more processors;
memory that stores computer-executable instructions that are executable by the one or more processors to cause the system to:
cause, via access to a data clean room, generation of a first privacy-compliant dataset associated with a first entity and a second privacy-compliant dataset associated with a second entity based at least in part on first data associated with the first entity and second data associated with the second entity, wherein the data clean room includes one or more permissions that restrict whether at least one of the first entity or the second entity may perform one or more operations using at least a portion of the first data or at least a portion of the second data via the data clean room;
host a first machine learning (ML) model and a second ML model, the first ML model trained using the first privacy-compliant dataset and the second ML model trained using the second privacy-compliant dataset, the first privacy-compliant dataset associated with a first entity that is distinct from a second entity associated with the second privacy compliant dataset or the first ML model comprises embedded proprietary parameters caused to be obfuscated by the first entity to prevent the second entity from accessing the proprietary parameters;
obtain a digital content request including a dataset comprising user data and digital content data;
process the dataset using the first ML model to generate a first score representing a first probability of a user engaging digital content;
process the dataset using the second ML model to generate a second score representing a second probability of a user engaging digital content; and
select at least one digital content from a plurality of digital content based on one or more of the first and second scores.
6 . The system of claim 5 , wherein the first ML model is provided by the first entity and the second ML model is provided by the second entity, the second entity providing a service to host the first ML model provided by the first entity, the service provided by the second entity further provided to receive the dataset and cause the first and second ML models to process the dataset.
7 . The system of claim 5 , wherein the user data and the digital content data comprise real-time data triggered when a digital content opportunity occurs on a computing device of the user, the user data comprising anonymized information about the user and the digital content data comprising digital content inventory information including placement type and size for the at least one digital content, the real-time data further comprising contextual information associated with media that will present the at least one digital content and device information comprising at least an operating system of the computing device.
8 . The system of claim 5 , wherein the memory that stores the computer executable instructions that are executable by the one or more processors are to further cause the system to:
store first anonymized data and second anonymized data, the first anonymized data isolated from the second anonymized data, and wherein the first privacy-compliant dataset includes at least data from the first anonymized data and the second privacy-compliant dataset includes at least data from the second anonymized data.
9 . The system of claim 5 , wherein selecting the at least one digital content from the plurality of digital content is based on the first and second scores.
10 . The system of claim 5 , wherein the first score is a first propensity score between 0 and 1 and the second score is a second propensity score between 0 and 1.
11 . The system of claim 5 , wherein the first ML model is based on an untrained ML model provided by an entity providing a service to select the at least one digital content, the untrained ML model comprising adjustable input features.
12 . The system of claim 5 , wherein the first and second ML models are associated with a programmatic digital content service comprising at least a real-time bidding engine to host the first and second ML models, the programmatic digital content service further comprising a data management layer to ingress the dataset comprising the user data and the digital content data.
13 . A non-transitory computer-readable storage medium storing thereon executable instructions that, as a result of being executed by one or more processors of a computer system, cause the computer system to:
cause, via access to a data clean room, generation of a first privacy-compliant dataset associated with a first entity and a second privacy-compliant dataset associated with a second entity based at least in part on first data associated with the first entity and second data associated with the second entity, wherein the data clean room includes one or more permissions that restrict whether at least one of the first entity or the second entity may perform one or more operations using at least a portion of the first data or at least a portion of the second data via the data clean room;
host a first machine learning (ML) model and a second ML model, the first ML model trained using the first privacy-compliant dataset and the second ML model trained using the second privacy-compliant dataset, the first privacy-compliant dataset associated with a first entity that is distinct from a second entity associated with the second privacy-compliant dataset or the first ML model comprises embedded proprietary parameters caused to be obfuscated by the first entity to prevent the second entity from accessing the proprietary parameters;
obtain a digital content request including a dataset comprising user data and digital content data;
process the dataset using the first ML model to generate a first score representing a first probability of a user engaging digital content;
process the dataset using the second ML model to generate a second score representing a second probability of a user engaging digital content; and
select at least one digital content from a plurality of digital content based on one or more of the first and second scores.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein the first ML model is provided by the first entity and the second ML model is provided by the second entity, the second entity providing a service to host the first ML model provided by the first entity, the service provided by the second entity further provided to receive the dataset and cause the first and second ML models to process the dataset.
15 . The non-transitory computer-readable storage medium of claim 13 , wherein the user data and the digital content data comprise real-time data triggered when a digital content opportunity occurs on a computing device of the user, the user data comprising anonymized information about the user and the digital content data comprising digital content inventory information including placement type and size for the at least one digital content, the real-time data further comprising contextual information associated with media that will present the at least one digital content and device information comprising at least an operating system of the computing device.
16 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, as a result of being executed by the one or more processors, further cause the computer system to:
store first anonymized data and second anonymized data, the first anonymized data isolated from the second anonymized data, and wherein the first privacy-compliant dataset includes at least data from the first anonymized data and the second privacy-compliant dataset includes at least data from the second anonymized data.
17 . The non-transitory computer-readable storage medium of claim 13 , wherein selecting the at least one digital content from the plurality of digital content is based the first and second scores.
18 . The non-transitory computer-readable storage medium of claim 13 , wherein the first score is a first propensity score between 0 and 1 and the second score is a second propensity score between 0 and 1.
19 . The non-transitory computer-readable storage medium of claim 13 , wherein the first ML model is based on an untrained ML model provided by an entity providing a service to select the at least one digital content, the untrained ML model comprising adjustable input features.
20 . The non-transitory computer-readable storage medium of claim 13 , wherein the first and second ML models are associated with a programmatic digital content service comprising at least a real-time bidding engine to host the first and second ML models, the programmatic digital content service further comprising a data management layer to ingress the dataset comprising the user data and the digital content data.