IP Library › Granted Patent US 12,105,837
Granted Patent B2
US 12,105,837 · App. 17/517,465 · Granted Oct 1, 2024

Generating private synthetic training data for training machine-learning models

Inventors: Christopher Lawrence LaTerza (Issaquah, WA); Girish Kumar (Santa Clara, CA); David Benjamin Levitan (Bothell, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F21/6245G06F18/214G06F40/40G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,105,837
App. No.
17/517,465
Granted
Oct 1, 2024
Kind
B2
Abstract

A method and system for generating synthetic privacy preserving training data for training a language classifier machine-learning (ML) model includes receiving a request to generate the synthetic privacy-preserving training data for the language classifier ML model, retrieving labeled training data associated with training the language classifier ML model, providing the labeled training data, one or more privacy parameters, and a domain type associated with the labeled training data to a synthetic data generation ML model, the synthetic data generation ML model being configured to generate synthetic training data in a privacy-persevering manner, receiving synthetic privacy-preserving training data as an output from the synthetic data generation ML model, and providing the synthetic privacy preserving training data to the language classifier ML model for training the language classifier ML model in classifying text.

Claims (55)

1. A data processing system comprising:

a processor; and

a memory in communication with the processor, the memory comprising executable instructions that, when executed by, the processor, cause the data processing system to perform functions of:

receiving a request to generate synthetic training data for a language classifier machine-learning (ML) model;

retrieving labeled training data associated with training the language classifier ML model;

providing the labeled training data, one or more privacy parameters, and a domain type associated with the labeled training data to a synthetic data generation ML model, the synthetic data generation ML model being configured to generate synthetic training data in a privacy-persevering manner,

wherein the privacy parameters include one or more differential privacy parameters having values that are dependent on a level of privacy required and the domain type;

receiving synthetic privacy-preserving training data as an output from the synthetic data generation ML model; and

providing the synthetic privacy-preserving training data to the language classifier ML model for training the language classifier ML model in classifying text,

wherein the synthetic privacy-preserving training data includes more training data than the retrieved labeled training data, and

wherein the synthetic privacy-preserving training data excludes private data in accordance with the privacy parameters and a predetermined leakage threshold for including private data.

2. The data processing system of claim 1 , wherein the instructions further cause the processor to cause the data processing system to perform functions of:

analyzing the synthetic privacy-preserving training data to determine that the synthetic privacy-preserving training data does not meet the leakage threshold for including private data; and

responsive to determining that the synthetic privacy-preserving training data does not meet the leakage threshold, removing one or more identified private data points from the synthetic privacy-preserving training data.

3. The data processing system of claim 2 , wherein the leakage threshold is a parameter associated with an acceptable amount of private data in the synthetic privacy-preserving training data.

4. The data processing system of claim 1 , wherein the instructions further cause the processor to cause the data processing system to perform functions of:

generating an initial prompt from one or more datasets of the labeled training data, and

providing the initial prompt as an input to the synthetic data generation ML model for generating the synthetic privacy-preserving training data.

5. The data processing system of claim 4 , wherein generating the initial prompt includes sampling one or more privacy-preserving words from one or more datapoints in the labeled training data.

6. The data processing system of claim 1 , wherein the synthetic data generation ML model is a generative adversarial network (GAN) model.

7. The data processing system of claim 1 , wherein the training data includes user feedback data and the language classifier ML model is used for classifying the user feedback data.

8. A method for synthetic privacy-preserving training data for training a language classifier machine-learning (ML) model, comprising:

receiving a request to generate the synthetic privacy-preserving training data for the language classifier ML model;

retrieving labeled training data associated with training the language classifier ML model;

providing the labeled training data, one or more privacy parameters, and a domain type associated with the labeled training data to a synthetic data generation ML model, the synthetic data generation ML model being configured to generate synthetic training data in a privacy-persevering manner,

wherein the privacy parameters include one or more differential privacy parameters having values that are dependent on a level of privacy required and the domain type;

receiving synthetic privacy-preserving training data as an output from the synthetic data generation ML model; and

providing the synthetic privacy-preserving training data to the language classifier ML model for training the language classifier ML model in classifying text,

wherein the synthetic privacy-preserving training data includes more training data than the retrieved labeled training data, and

wherein the synthetic privacy-preserving training data excludes private data in accordance with the privacy parameters and a predetermined leakage threshold for including private data.

9. The method of claim 8 , further comprising: analyzing the synthetic privacy-preserving training data to determine that the synthetic privacy-preserving training data does not meet the leakage threshold for including private data; and

responsive to determining that the synthetic privacy-preserving training data does not meet the leakage threshold, removing one or more identified private data points from the synthetic privacy-preserving training data.

10. The method of claim 9 , wherein the leakage threshold is a parameter associated with an acceptable amount of private data in the synthetic privacy-preserving training data.

11. The method of claim 8 , further comprising:

generating an initial prompt from one or more datasets of the labeled training data, and

providing the initial prompt as an input to the synthetic data generation ML model for generating the synthetic privacy-preserving training data.

12. The method of claim 11 , wherein generating the initial prompt includes sampling one or more privacy-preserving words from one or more datapoints in the labeled training data.

13. The method of claim 8 , wherein the synthetic data generation ML model is a generative adversarial network (GAN) model.

14. The method of claim 8 , wherein the training data includes user feedback data and the language classifier ML model is used for classifying the user feedback data.

15. A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to perform functions of:

receiving a request to generate synthetic training data for a language classifier machine-learning (ML) model;

retrieving labeled training data associated with training the language classifier ML model;

providing the labeled training data, one or more privacy parameters, and a domain type associated with the labeled training data to a synthetic data generation ML model, the synthetic data generation ML model being configured to generate synthetic training data in a privacy-persevering manner,

wherein the privacy parameters include one or more differential privacy parameters having values that are dependent on a level of privacy required and the domain type;

receiving synthetic privacy-preserving training data as an output from the synthetic data generation ML model; and

providing the synthetic privacy-preserving training data to the language classifier ML model for training the language classifier ML model in classifying text,

wherein the synthetic privacy-preserving training data includes more training data than the retrieved labeled training data, and

wherein the synthetic privacy-preserving training data excludes private data in accordance with the privacy parameters and a predetermined leakage threshold for including private data.

16. The non-transitory computer readable medium of claim 15 , wherein the instructions further cause the programmable device to perform functions of:

analyzing the synthetic privacy-preserving training data to determine that the synthetic privacy-preserving training data does not meet the leakage threshold for including private data; and

responsive to determining that the synthetic privacy-preserving training data does not meet the leakage threshold, removing one or more identified private data points from the synthetic privacy-preserving training data.

17. The non-transitory computer readable medium of claim 15 , wherein the instructions further cause the programmable device to perform functions of:

generating an initial prompt from one or more datasets of the labeled training data, and

providing the initial prompt as an input to the synthetic data generation ML model for generating the synthetic privacy-preserving training data.

18. The non-transitory computer readable medium of claim 15 , wherein the training data includes user feedback data and the language classifier ML model is used for classifying the user feedback data.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE FIRST ASSIGNOR'S NAME AND EXECUTION DATE INSIDE THE ASSIGNMENT DOCUMENT AND ON THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 057999 FRAME: 0215. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 24, 2022
From: LATERZA, CHRISTOPHER LAWRENCE; KUMAR, GIRISH; LEVITAN, DAVID BENJAMIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061783/0111 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: LATERZA, CHRISTOPHER; KUMAR, GIRISH; LEVITAN, DAVID BENJAMIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057999/0215 →
Continuity (1)
Related Publication 20230137378A1 · May 4, 2023
Cited By (1)
US 12,307,499