IP Library Granted Patent US 11,640,527
Granted Patent B2
US 11,640,527 · App. 16/658,399 · Granted May 2, 2023

Near-zero-cost differentially private deep learning with teacher ensembles

Inventors: Lichao Sun (Chicago, IL); Jia Li (Mountain View, CA); Caiming Xiong (Menlo Park, CA); Yingbo Zhou (San Jose, CA)
Assignee: Salesforce.com, Inc.
G06N3/08G06N3/04G06N3/0454G06N3/082G06T2207/00G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,527
App. No.
16/658,399
Granted
May 2, 2023
Kind
B2
Abstract

Systems and methods are provided for near-zero-cost (NZC) query framework or approach for differentially private deep learning. To protect the privacy of training data during learning, the near-zero-cost query framework transfers knowledge from an ensemble of teacher models trained on partitions of the data to a student model. Privacy guarantees may be understood intuitively and expressed rigorously in terms of differential privacy. Other features are also provided.

Claims (40)

1. A system for private deep learning, comprising:

a communication interface that receives model is used to process input data including sensitive information;

a memory storing a machine learning model and a plurality of processor-executable instructions; and

one or more processors that execute the plurality of processor-executable instructions,

wherein during a training process of the machine learning model, to:

partition a dataset including the sensitive information into a plurality of data splits;

train each of a plurality of teacher models in an ensemble with a respective one of the data splits;

in response to a query, generate, by each trained teacher model, a respective prediction;

aggregate the predictions generated by the teacher models in the ensemble into a label count vector by computing each entry in the label count vector based on a sum of predicted probabilities from the plurality of teacher models corresponding to a respective label;

perturb the label count vector for the teacher ensemble with noise and by adding a constant to a highest counted class vote in the label count vector; and

train a student model using the perturbed label count vector for the teacher ensemble.

2. The system of claim 1 , wherein aggregating the predictions is responsive to a query from the student model.

3. The system of claim 1 , wherein the noise used to perturb the label count vector is random.

4. The system of claim 1 , wherein the machine learning model is configured to use a query function.

5. The system of claim 1 , wherein during a training process, the machine learning model performs a noisy ArgMax operation on the label count vector for the teacher ensemble.

6. The system of claim 1 , wherein during a training process, the machine learning model performs an immutable noisy ArgMax operation on the label count vector for the teacher ensemble.

7. A method for training a machine learning model with private deep learning, comprising:

partitioning a dataset including the sensitive information into a plurality of data splits;

training each of a plurality of teacher models in an ensemble with a respective one of the data splits;

in response to a query, generating, by each trained teacher model, a respective prediction;

aggregating the predictions generated by the teacher models in the ensemble into a label count vector by computing each entry in the label count vector based on a sum of predicted probabilities from the plurality of teacher models corresponding to a respective label;

perturbing the label count vector for the teacher ensemble with noise and by adding a constant to a highest counted class vote in the label count vector; and

training a student model using the perturbed label count vector for the teacher ensemble.

8. The method of claim 7 , wherein aggregating the predictions is responsive to a query from the student model.

9. The method of claim 7 , wherein the noise used to perturb the label count vector is random.

10. The method of claim 7 , comprising the student model making a query function to the label count vector.

11. The method of claim 7 , comprising performing a noisy ArgMax operation on the label count vector for the teacher ensemble.

12. The method of claim 7 , comprising performing an immutable noisy ArgMax operation on the label count vector for the teacher ensemble.

13. A non-transitory machine-readable medium comprising a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method for training a machine learning model with private deep learning comprising:

partitioning a dataset including the sensitive information into a plurality of data splits;

training each of a plurality of teacher models in an ensemble with a respective one of the data splits;

in response to a query, generating, by each trained teacher model, a respective prediction;

aggregating the predictions generated by the teacher models in the ensemble into a label count vector by computing each entry in the label count vector based on a sum of predicted probabilities from the plurality of teacher models corresponding to a respective label;

perturbing the label count vector for the teacher ensemble with noise and by adding a constant to a highest counted class vote in the label count vector; and

training a student model using the perturbed label count vector for the teacher ensemble.

14. The non-transitory machine-readable medium of claim 13 , wherein aggregating the predictions is responsive to a query from the student model.

15. The non-transitory machine-readable medium of claim 13 , wherein the noise used to perturb the label count vector is random.

16. The non-transitory machine-readable medium of claim 13 , comprising the student model making a query function to the label count vector.

17. The non-transitory machine-readable medium of claim 13 , comprising performing a noisy ArgMax operation on the label count vector for the teacher ensemble.

18. The non-transitory machine-readable medium of claim 13 , comprising performing an immutable noisy ArgMax operation on the label count vector for the teacher ensemble.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0452 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2019
From: SUN, LICHAO; LI, JIA; XIONG, CAIMING; ZHOU, YINGBO
To: SALESFORCE.COM, INC.
Reel/Frame 050838/0740 →
Continuity (2)
Provisional Application 62906020 · Sep 25, 2019
Related Publication 20210089882A1 · Mar 25, 2021