IP Library › Granted Patent US 11,841,977
Granted Patent B2
US 11,841,977 · App. 17/173,628 · Granted Dec 12, 2023

Training anonymized machine learning models via generalized data generated using received trained machine learning models

Inventors: Abigail Goldsteen (Haifa, IL); Ariel Farkash (Shimshit, IL); Micha Gideon Moffie (Zichron Yaakov, IL); Gilad Ezov (Nesher, IL); Ron Shmelkin (Haifa, IL)
Assignee: International Business Machines Corporation
G06F21/6254G06F18/2148G06F18/2155G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,841,977
App. No.
17/173,628
Granted
Dec 12, 2023
Kind
B2
Abstract

An example system includes a processor to receive training data and predictions on the training data of a trained machine learning model to be anonymized. The processor is to generate generalized data from training data based on the predictions of the trained machine learning model on the training data. The processor is to train an anonymized machine learning model using the generalized data.

Claims (31)

1. A system, comprising a processor to:

receive training data and predictions on the training data of a trained machine learning model to be anonymized;

generate generalized data based on the predictions of the trained machine learning model on the training data; and

train an anonymized machine learning model using the generalized data.

2. The system of claim 1 , wherein the processor is to generate an anonymizer model based on the predictions of the trained machine learning model on the training data and generate the generalized data from the training data using the anonymizer model.

3. The system of claim 1 , wherein the processor is to use the predictions of the trained machine learning model on the training data as input to generate groups of similar records based on similar outputs from the trained machine learning model and generalize the groups to generate an anonymizer model used to generate the generalized data.

4. The system of claim 1 , wherein the generalized data comprises representative values in the same domain as original features used to train the anonymized machine learning model.

5. The system of claim 1 , wherein the training data comprises unlabeled data and the processor is to label the unlabeled data based on the predictions of the trained machine learning model.

6. The system of claim 1 , wherein the predictions comprise outputs from a layer of the trained machine learning model that is prior to a final classification layer.

7. The system of claim 1 , wherein the trained machine learning model comprises a complex model, and the anonymized machine learning model comprises an anonymized part of the complex model.

8. A computer-implemented method, comprising:

receiving, via a processor, training data and predictions on the training data of a trained machine learning model to be anonymized;

generating, via the processor, an anonymizer model based on the predictions of the trained machine learning model on the training data;

anonymizing, via the processor, the training data via the anonymizer model to generate generalized data; and

retraining, via the processor, the trained machine learning model using the generalized data to generate an anonymized machine learning model.

9. The computer-implemented method of claim 8 , wherein generating the anonymizer model comprises training a decision tree using predictions of the trained machine learning model.

10. The computer-implemented method of claim 8 , wherein generating the anonymizer model comprises using a two phase clustering algorithm comprising a coarse clustering phase and a sub-clustering phase.

11. The computer-implemented method of claim 8 , wherein generating the anonymizer model comprises using the predictions of the trained machine learning model on the training data as input to generate groups of similar records and generalizing the groups to generate the anonymizer model.

12. The computer-implemented method of claim 8 , wherein anonymizing the training data comprises replacing data points in each cluster or bucket of similar inputs with a representative value for the cluster or the bucket.

13. The computer-implemented method of claim 8 , comprising receiving the trained machine learning model and generating the predictions, via the trained learning model, on the training data.

14. The computer-implemented method of claim 8 , wherein retraining the trained machine learning model comprises retraining parts of the trained machine learning model, wherein the trained machine learning model comprises a complex model.

15. The computer-implemented method of claim 8 , where the training data that is generalized for creating the anonymized machine learning model is different from the training data used to train the trained machine learning model.

16. A computer program product for anonymizing machine learning models, the computer program product comprising a computer-readable storage medium having program code embodied therewith, wherein the computer-readable storage medium is not a transitory signal per se, the program code executable by a processor to cause the processor to:

receive training data and predictions on the training data of a trained machine learning model to be anonymized;

generate an anonymizer model based on the predictions of the trained machine learning model on the training data;

anonymize the training data via the anonymizer model to generate generalized data; and

retrain the trained machine learning model using the generalized data to generate an anonymized machine learning model.

17. The computer program product of claim 16 , further comprising program code executable by the processor to train a decision tree using predictions of the trained machine learning model.

18. The computer program product of claim 16 , further comprising program code executable by the processor to execute a two phase clustering comprising a coarse clustering phase and a sub-clustering phase.

19. The computer program product of claim 16 , further comprising program code executable by the processor to use the predictions of the trained machine learning model on the training data as input to generate groups of similar records and generalize the groups to generate the anonymizer model.

20. The computer program product of claim 16 , further comprising program code executable by the processor to replace data points in each cluster or bucket of similar inputs with a representative value for the cluster or the bucket.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: GOLDSTEEN, ABIGAIL; FARKASH, ARIEL; MOFFIE, MICHA GIDEON; EZOV, GILAD; SHMELKIN, RON
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055236/0069 →
Continuity (1)
Related Publication 20220253554A1 · Aug 11, 2022