IP Library Granted Patent US 12,373,729
Granted Patent B2
US 12,373,729 · App. 17/085,699 · Granted Jul 29, 2025

System and method for federated learning with local differential privacy

Inventors: Jianwei Qian (Mountain View, CA); Lichao Sun (Easton, PA); Xun Chen (Fremont, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N20/00G06F21/6263
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,729
App. No.
17/085,699
Granted
Jul 29, 2025
Kind
B2
Abstract

In one embodiment, a method includes accessing a plurality of initial gradients associated with a machine-learning model from a data store associated with a first electronic device, selecting one or more of the plurality of initial gradients for perturbation, generating one or more perturbed gradients for the one or more selected initial gradients based on a gradient-perturbation model, respectively, wherein for each selected initial gradient: an input to the gradient-perturbation model comprises the selected initial gradient having a value x, the gradient-perturbation model changes x into a first continuous value with a first probability or a second continuous value with a second probability, and the first and second probabilities are determined based on x, and sending the one or more perturbed gradients from the first electronic device to a second electronic device.

Claims (408)

1. A method comprising, by one or more processors of a first electronic device:

by one or more of the processors, accessing, from a data store associated with the first electronic device, a plurality of initial gradients that determine the output of a machine-learning model for a given input;

by one or more of the processors, selecting one or more of the plurality of initial gradients for perturbation;

by one or more of the processors, obfuscating user-identifying information of a user corresponding to the plurality of initial gradients by generating, based on a gradient-perturbation model, one or more perturbed gradients for the one or more selected initial gradients, respectively, wherein for each selected initial gradient:

an input to the gradient-perturbation model comprises the selected initial gradient having a value x,

the gradient-perturbation model perturbs the value x into either of two endpoints of a perturbation range, the two endpoints comprising (1) a first continuous value,

comprising a lower boundary of the perturbation range defined by a center value c and a distance r from the center c, with a first probability and (2) a second continuous value, comprising an upper boundary of the perturbation range, with a second probability, wherein (1) the first and second probabilities are determined based on x, (2) the first and second probabilities sum to 1, and (3) the first and second probabilities preserve accuracy when averaging initial gradients by making the expectation value of each particular perturbed gradient, E (A (x)), equal to its corresponding initial gradient x, even though the particular perturbed gradient is not equal to the corresponding initial gradient; and

by one or more of the processors, sending, from the first electronic device to a second electronic device, the one or more perturbed gradients.

2. The method of claim 1 , further comprising:

determining, based on one or more privacy policies, that one or more of the plurality of initial gradients should be perturbed.

3. The method of claim 1 , further comprising:

receiving, at the first electronic device from the second electronic device, a plurality of weights of the machine-learning model, wherein the plurality of weights are determined based on the one or more perturbed gradients; and

determining, by the first electronic device, a plurality of new gradients for the plurality of weights.

4. The method of claim 1 , wherein the perturbation of the one or more selected gradients is performed according to:

A

(

x

)

=

{

c

+

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

c

-

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

-

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

,

wherein:

A (x) represents the perturbed value of x,

c represents a center value of the perturbation range,

r represents a distance from the center value to boundaries of the perturbation range,

each selected initial gradient is clipped into the value range, and

∈ is a positive real number determined based on a local differential policy.

5. The method of claim 1 , further comprising:

shuffling the one or more perturbed gradients to a random order;

wherein the one or more perturbed gradients are sent based on the random order.

6. A computer-readable non-transitory storage media comprising instructions executable by a processor to:

access, from a data store associated with the first electronic device, a plurality of initial gradients that determine the output of a machine-learning model for a given input;

select one or more of the plurality of initial gradients for perturbation;

obfuscate user-identifying information of a user corresponding to the plurality of initial gradients by generating, based on a gradient-perturbation model, one or more perturbed gradients for the one or more selected initial gradients, respectively, wherein for each selected initial gradient:

an input to the gradient-perturbation model comprises the selected initial gradient having a value x,

the gradient-perturbation model perturbs the value x into either of two endpoints of a perturbation range, the two endpoints comprising (1) a first continuous value, comprising a lower boundary of the perturbation range defined by a center value c and a distance r from the center c, with a first probability and (2) a second continuous value, comprising an upper boundary of the perturbation range, with a second probability, wherein (1) the first and second probabilities are determined based on x, (2) the first and second probabilities sum to 1, and (3) the first and second probabilities preserve accuracy when averaging initial gradients by making the expectation value of each particular perturbed gradient, E (A (x)), equal to its corresponding initial gradient x, even though the particular perturbed gradient is not equal to the corresponding initial gradient; and

send, from the first electronic device to a second electronic device, the one or more perturbed gradients.

7. The media of claim 6 , wherein the instructions are further executable by the processor to:

determine, based on one or more privacy policies, that one or more of the plurality of initial gradients should be perturbed.

8. The media of claim 6 , wherein the instructions are further executable by the processor to:

receive, at the first electronic device from the second electronic device, a plurality of weights of the machine-learning model, wherein the plurality of weights are determined based on the one or more perturbed gradients; and

determine, by the first electronic device, a plurality of new gradients for the plurality of weights.

9. The media of claim 6 , wherein the perturbation of the one or more selected gradients is performed according to:

A

(

x

)

=

{

c

+

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

c

-

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

-

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

,

wherein:

A (x) represents the perturbed value of x,

c represents a center value of the perturbation range,

r represents a distance from the center value to boundaries of the perturbation range,

each selected initial gradient is clipped into the value range, and

∈ is a positive real number determined based on a local differential policy.

10. The media of claim 6 , wherein the instructions are further executable by the processor to:

shuffle the one or more perturbed gradients to a random order;

wherein the one or more perturbed gradients are sent based on the random order.

11. A system comprising:

one or more processors; and

a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

access, from a data store associated with the first electronic device, a plurality of initial gradients that determine the output of a machine-learning model for a given input;

select one or more of the plurality of initial gradients for perturbation;

obfuscate user-identifying information of a user corresponding to the plurality of initial gradients by generating, based on a gradient-perturbation model, one or more perturbed gradients for the one or more selected initial gradients, respectively, wherein for each selected initial gradient:

an input to the gradient-perturbation model comprises the selected initial gradient having a value x,

the gradient-perturbation model perturbs the value x into either of two endpoints of a perturbation range, the two endpoints comprising (1) a first continuous value, comprising a lower boundary of the perturbation range defined by a center value c and a distance r from the center c, with a first probability and (2) a second continuous value, comprising an upper boundary of the perturbation range, with a second probability, wherein (1) the first and second probabilities are determined based on x, (2) the first and second probabilities sum to 1, and (3) the first and second probabilities preserve accuracy when averaging initial gradients by making the expectation value of each particular perturbed gradient, E(A(x)), equal to its corresponding initial gradient x, even though the particular perturbed gradient is not equal to the corresponding initial gradient; and

send, from the first electronic device to a second electronic device, the one or more perturbed gradients.

12. The system of claim 11 , wherein the processors are further operable when executing the instructions to:

determine, based on one or more privacy policies, that one or more of the plurality of initial gradients should be perturbed.

13. The system of claim 11 , wherein the processors are further operable when executing the instructions to:

receive, at the first electronic device from the second electronic device, a plurality of weights of the machine-learning model, wherein the plurality of weights are determined based on the one or more perturbed gradients; and

determine, by the first electronic device, a plurality of new gradients for the plurality of weights.

14. The system of claim 11 , wherein the perturbation of the one or more selected gradients is performed according to:

A

(

x

)

=

{

c

+

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

c

-

r

·

e

ϵ

+

1

e

ϵ

-

1

,

with

probability

-

(

x

-

c

)

(

e

ϵ

-

1

)

+

r

(

e

ϵ

+

1

)

2

r

(

e

ϵ

+

1

)

,

wherein:

A (x) represents the perturbed value of x,

c represents a center value of the perturbation range,

r represents a distance from the center value to boundaries of the perturbation range,

each selected initial gradient is clipped into the value range, and

∈ is a positive real number determined based on a local differential policy.

15. The system of claim 11 , wherein the processors are further operable when executing the instructions to:

shuffle the one or more perturbed gradients to a random order;

wherein the one or more perturbed gradients are sent based on the random order.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2020
From: QIAN, JIANWEI; SUN, LICHAO; CHEN, XUN
To: SAMSUNG ELECTRONICS COMPANY, LTD.
Reel/Frame 054227/0843 →
Continuity (2)
Provisional Application 63031531 · May 28, 2020
Related Publication 20210374605A1 · Dec 2, 2021
References Cited (15)
US 20110139748A1 · Donnelly · 2011 [cited by examiner]
US 20120212375A1 · Depree, IV · 2012 [cited by examiner]
US 20170134434A1 · Allen · 2017 [cited by examiner]
US 20190220703A1 · Prakash · 2019 [cited by applicant]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
KR 20190103090 · 2019 [cited by applicant]
KR 20200010480 · 2020 [cited by applicant]
Shokri, Reza, and Vitaly Shmatikov. “Privacy-preserving deep learning.” In [cited by applicant]
Jayaraman, Bargav, Lingxiao Wang, David Evans, and Quanquan Gu. “Distributed learning without distress: Privacy-preserving empirical risk minimization.” In [cited by applicant]
Bebensee, Bjorn. “Local differential privacy: a tutorial.” [cited by applicant]
Kairouz, Peter, Sewoong Oh, and Pramod Viswanath. “Extremal mechanisms for local differential privacy.” In [cited by applicant]
Erlingsson, Ulfar, Vasyl Pihur, and Aleksandra Korolova. “Rappor: Randomized aggregatable privacy-preserving ordinal response.” In Proceedings of the 2014 ACM SIGSAC conference on computer and communications security, p… [cited by applicant]
Dwork, Cynthia. “Differential privacy: A survey of results.” In [cited by applicant]
Truex, Stacey, Ling Liu, Ka-Ho Chow, Mehmet Emre Gursoy, and Wenqi Wei. “LDP-Fed: federated learning with local differential privacy.” In [cited by applicant]
Seif, Mohamed, Ravi Tandon, and Ming Li. “Wireless federated learning with local differential privacy.” arXiv preprint arXiv:2002.05151, pp. 1-13, Feb. 13, 2020. [cited by applicant]