Method and server for defending service from personal privacy inference attack
View Patent ↗A computer-implemented method for preventing leaking a personal privacy from a chatbot under black-box personal attribute inference attack is provided. The chatbot is provided via a neural network executed by a processor of a server. The method includes: training, by the processor, a Language Model (LM) of the chatbot according to utility objectives; applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune a target LM of the chatbot by using a fake attacker model and pre-define attributes with annotated datasets; and using, by the processor, the target LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by external predictor and a security level of the chatbot is assured.
1 . A computer-implemented method for preventing leaking a personal privacy from a chatbot under black-box persona inference attack, wherein the chatbot is provided via a neural network executed by a processor of a server, the method comprises:
training, by the processor, a Language Model (LM) of the chatbot according to utility objectives;
applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune the LM of the chatbot by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the LM in place of real attacker model details for updating the LM during actual attacks; and
using, by the processor, the fine-tuned LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by an external predictor and a security level of the chatbot is assured;
wherein the utility objectives comprising a LM loss; and
wherein the LM loss is an objective function of the LM, and the objective function is presented by a formula below:
L
f
(
U
;
θ
f
)
=
-
∑
i
=
1
❘
"\[LeftBracketingBar]"
U
❘
"\[RightBracketingBar]"
log
(
Pr
(
w
i
❘
"\[LeftBracketingBar]"
c
,
w
0
,
w
1
,
…
,
w
i
-
1
)
)
where L f refers to a loss function of the LM model; f refers to the LM model; θ f refers to parameters of the LM; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the LM f with a given utterance U={w 0 , w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;
wherein the LM of the chatbot is fine-tuned by:
sending, by a client terminal of a user, a first query data to the chatbot;
obtaining by the fake attacker model, a first content of the first query data,
generating by the LM of the chatbot, a response data having a response content corresponding to the content of the query data;
obtaining, by the fake attacker model, the response content;
sending, by the client terminal a second query data to the chatbot;
obtaining, by the fake attacker model, a second content of the second query data; and
outputting, by the fake model, a prediction data having a predicted persona of the user according to the first content, the response content and the second content.
2 . The method of claim 1 , wherein the defense objectives comprising one or a combination of following:
a KL (Kullback-Leibler) loss; and
a MI (Mutual Information) loss.
3 . The method of claim 2 , wherein an objective function of the KL loss is presented by a formula below:
L
kl
(
u
;
,
θ
f
)
=
-
1
C
∑
k
=
0
C
-
1
Pr
(
k
❘
"\[LeftBracketingBar]"
f
(
u
)
,
)
where L kl refers to a loss function of the KL loss; refers to parameters of the fake attacker; u refers to utterance; k refers to personal attribute label index; C refers to a total number of predefined personal attributes; and f(u) refers to hidden states of the chatbot.
4 . The method of claim 2 , wherein an objective function of the MI loss is presented by a formula below:
min
θ
f
max
Ψ
𝔼
q
(
f
(
u
)
)
[
log
p
Ψ
(
s
❘
"\[LeftBracketingBar]"
f
(
u
)
)
]
where E q refers to a loss function of the KL loss; p Ψ (s|f(u)) refers to a distribution function used to approximate q(s|f(u)) which refers to probability distribution of model f parameterized by θ f , Ψ refers to the attacker model that manages to infer s from f(u).
5 . The method of claim 1 , wherein the fake attacker comprises:
a projection layer comprising a plurality of fully connected layers; and
a softmax activation function layer.
6 . The method of claim 5 , wherein a loss function of the fake attacker is presented by a formula below:
( u kj ,s kj ; )= CE ( p ( f ( u kj ), s kj )
where is the loss function of the fake attacker; u kj represents j-th utterance from k-th conversation; CE refers to cross-entropy between personal label s kj and personal attribute predictor's output p (f(u kj )); and represents parameters of a persona predictor model of the fake attacker.
7 . A server for preventing leaking a personal privacy from a chatbot under black-box personal attribute inference attack, wherein the chatbot is provided via a neural network executed by a processor of the server, comprising:
the processor, configured to execute machine instructions to implement a computer-implemented method, the method comprising:
training, by the processor, a Language Model (LM) of the chatbot according to utility objectives;
applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune the LM of the chatbot by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the LM in place of real attacker model details for updating the LM during actual attacks; and
using, by the processor, the fine-tuned LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by an external predictor; and a security level of the chatbot is assured;
wherein the utility objectives comprising a LM loss; and
wherein the LM loss is an objective function of the LM, and the objective function is presented by a formula below:
L
f
(
U
;
θ
f
)
=
-
∑
i
=
1
❘
"\[LeftBracketingBar]"
U
❘
"\[RightBracketingBar]"
log
(
Pr
(
w
i
❘
"\[LeftBracketingBar]"
c
,
w
0
,
w
1
,
…
,
w
i
-
1
)
)
where L f refers to a loss function of the LM model; f refers to the LM model; θ f refers to parameters of the LM; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the LM f with a given utterance U={w 0 ,w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;
wherein the LM of the chatbot is fine-tuned by:
sending, by a client terminal of a user, a first query data to the chatbot;
obtaining, by the fake attacker model, a first content of the first query data;
generating, by the IM of the chatbot, a response data having a response content corresponding to the content of the query data;
obtaining by the fake attacker model, the response content;
sending by the client terminal, a second query data to the chatbot;
obtaining, by the fake attacker model, a second content of the second query data; and
outputting by the fake attacker model, a prediction data having a predicted persona of the user according to the first content the response content and the second content.
8 . The server of claim 7 , wherein the defense objectives comprising one or a combination of following:
a KL (Kullback-Leibler) loss; and
a MI (Mutual Information) loss.
9 . The server of claim 8 , wherein an objective function of the KL loss is presented by a formula below:
L
kl
(
u
;
,
θ
f
)
=
-
1
C
∑
k
=
0
C
-
1
Pr
(
k
❘
"\[LeftBracketingBar]"
f
(
u
)
,
)
where L kl refers to a loss function of the KL loss; refers to parameters of the fake attacker; u refers to utterance; k refers to personal attribute label index; C refers to a total number of predefined personal attributes; and f(u) refers to hidden states of the chatbot.
10 . The server of claim 8 , wherein an objective function of the MI loss is presented by a formula below:
min
θ
f
max
Ψ
𝔼
q
(
f
(
u
)
)
[
log
p
Ψ
(
s
❘
"\[LeftBracketingBar]"
f
(
u
)
)
]
where E q refers to a loss function of the KL loss; p Ψ (s|f(u)) refers to a distribution function used to approximate q(s|f(u)) which refers to probability distribution of model f parameterized by θ f ; Ψ refers to the attacker model that manages to infer s from f(u).
11 . The server of claim 7 , wherein the fake attacker comprises:
a projection layer comprising a plurality of fully connected layers; and
a softmax activation function layer.
12 . The server of claim 11 , wherein a loss function of the fake attacker is presented by a formula below:
( u kj ,s kj ; )= CE ( p ( f ( u kj ), s kj )
where is the loss function of the fake attacker; u kj represents j-th utterance from k-th conversation; CE refers to cross-entropy between personal attribute label s kj and personal attribute predictor's output p (f(u kj )); and represents parameters of a persona predictor model p of the fake attacker.
13 . A computer-implemented method for preventing leaking a personal privacy from a service under personal privacy inference attack, wherein the service is provided via a neural network executed by a processor of a server, the method comprises:
training, by the processor, a main algorithm model of the service according to utility objectives;
applying, by the processor, one or more defense objectives with attribute predictor to fine-tune the main algorithm model of the service by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the main algorithm model in place of the real attacker model details for updating the main algorithm model during actual attacks; and
using, by the processor, the ta fine-tuned main algorithm model on the service to defend inference attack, such that the personal privacy of content inputted and sent to the service cannot be predicted by an external predictor; and a security level of the service is assured;
wherein the utility objectives comprising a LM loss; and
wherein the LM loss is an objective function of the main algorithm model, and the objective function is presented by a formula below:
L
f
(
U
;
θ
f
)
=
-
∑
i
=
1
|
U
|
log
(
P
r
(
w
i
|
c
,
w
0
,
w
1
,
…
,
w
i
-
1
)
)
where L f refers to a loss function of the main algorithm model; f refers to the main algorithm model; θ f refers to parameters of the main algorithm model; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the main algorithm model f with a given utterance U={w 0 , w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;
wherein the main algorithm model of the chatbot is fine tuned by:
sending, by a client terminal of a user, a first query data to the chatbot;
obtaining, by the fake attacker model, a first content of the first query data;
generating, by the main algorithm model of the chatbot, a response data having a response content corresponding to the content of the query data;
obtaining by the fake attacker model, the response content;
sending, by the client terminal, a second query data to the chatbot;
obtaining, by the fake attacker model a second content of the second query data, and
outputting by the fake attacker model, a prediction data having a predicted persona of the user according to the first content, the response content and the second content.