IP Library › Granted Patent US 12,651,086
Granted Patent B2
US 12,651,086 · App. 17/985,147 · Granted Jun 9, 2026

Method and server for defending service from personal privacy inference attack

Inventors: Yangqiu Song (Hong Kong, CN); Haoran Li (Hong Kong, CN)
Assignee: The Hong Kong University of Science and Technology
G06F21/6245G06N3/094
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,086
App. No.
17/985,147
Granted
Jun 9, 2026
Kind
B2
Abstract

A computer-implemented method for preventing leaking a personal privacy from a chatbot under black-box personal attribute inference attack is provided. The chatbot is provided via a neural network executed by a processor of a server. The method includes: training, by the processor, a Language Model (LM) of the chatbot according to utility objectives; applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune a target LM of the chatbot by using a fake attacker model and pre-define attributes with annotated datasets; and using, by the processor, the target LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by external predictor and a security level of the chatbot is assured.

Claims (344)

1 . A computer-implemented method for preventing leaking a personal privacy from a chatbot under black-box persona inference attack, wherein the chatbot is provided via a neural network executed by a processor of a server, the method comprises:

training, by the processor, a Language Model (LM) of the chatbot according to utility objectives;

applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune the LM of the chatbot by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the LM in place of real attacker model details for updating the LM during actual attacks; and

using, by the processor, the fine-tuned LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by an external predictor and a security level of the chatbot is assured;

wherein the utility objectives comprising a LM loss; and

wherein the LM loss is an objective function of the LM, and the objective function is presented by a formula below:

L

f

(

U

;

θ

f

)

=

-

∑

i

=

1

❘

"\[LeftBracketingBar]"

U

❘

"\[RightBracketingBar]"

log

⁡

(

Pr

⁡

(

w

i

⁢

❘

"\[LeftBracketingBar]"

c

,

w

0

,

w

1

,

…

,

w

i

-

1

)

)

where L f refers to a loss function of the LM model; f refers to the LM model; θ f refers to parameters of the LM; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the LM f with a given utterance U={w 0 , w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;

wherein the LM of the chatbot is fine-tuned by:

sending, by a client terminal of a user, a first query data to the chatbot;

obtaining by the fake attacker model, a first content of the first query data,

generating by the LM of the chatbot, a response data having a response content corresponding to the content of the query data;

obtaining, by the fake attacker model, the response content;

sending, by the client terminal a second query data to the chatbot;

obtaining, by the fake attacker model, a second content of the second query data; and

outputting, by the fake model, a prediction data having a predicted persona of the user according to the first content, the response content and the second content.

2 . The method of claim 1 , wherein the defense objectives comprising one or a combination of following:

a KL (Kullback-Leibler) loss; and

a MI (Mutual Information) loss.

3 . The method of claim 2 , wherein an objective function of the KL loss is presented by a formula below:

L

kl

(

u

;

,

θ

f

)

=

-

1

C

⁢

∑

k

=

0

C

-

1

Pr

⁡

(

k

⁢

❘

"\[LeftBracketingBar]"

f

⁡

(

u

)

,

)

where L kl refers to a loss function of the KL loss; refers to parameters of the fake attacker; u refers to utterance; k refers to personal attribute label index; C refers to a total number of predefined personal attributes; and f(u) refers to hidden states of the chatbot.

4 . The method of claim 2 , wherein an objective function of the MI loss is presented by a formula below:

min

θ

f

max

Ψ

⁢

𝔼

q

⁡

(

f

⁡

(

u

)

)

[

log

⁢

p

Ψ

(

s

⁢

❘

"\[LeftBracketingBar]"

f

⁡

(

u

)

)

]

where E q refers to a loss function of the KL loss; p Ψ (s|f(u)) refers to a distribution function used to approximate q(s|f(u)) which refers to probability distribution of model f parameterized by θ f , Ψ refers to the attacker model that manages to infer s from f(u).

5 . The method of claim 1 , wherein the fake attacker comprises:

a projection layer comprising a plurality of fully connected layers; and

a softmax activation function layer.

6 . The method of claim 5 , wherein a loss function of the fake attacker is presented by a formula below:

( u kj ,s kj ; )= CE ( p ( f ( u kj ), s kj )

where is the loss function of the fake attacker; u kj represents j-th utterance from k-th conversation; CE refers to cross-entropy between personal label s kj and personal attribute predictor's output p (f(u kj )); and represents parameters of a persona predictor model of the fake attacker.

7 . A server for preventing leaking a personal privacy from a chatbot under black-box personal attribute inference attack, wherein the chatbot is provided via a neural network executed by a processor of the server, comprising:

the processor, configured to execute machine instructions to implement a computer-implemented method, the method comprising:

training, by the processor, a Language Model (LM) of the chatbot according to utility objectives;

applying, by the processor, one or more defense objectives with personal attribute predictor to fine-tune the LM of the chatbot by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the LM in place of real attacker model details for updating the LM during actual attacks; and

using, by the processor, the fine-tuned LM on the chatbot to defend inference attack, such that the personal privacy of content inputted and sent to the chatbot cannot be predicted by an external predictor; and a security level of the chatbot is assured;

wherein the utility objectives comprising a LM loss; and

wherein the LM loss is an objective function of the LM, and the objective function is presented by a formula below:

L

f

(

U

;

θ

f

)

=

-

∑

i

=

1

❘

"\[LeftBracketingBar]"

U

❘

"\[RightBracketingBar]"

log

⁡

(

Pr

⁡

(

w

i

⁢

❘

"\[LeftBracketingBar]"

c

,

w

0

,

w

1

,

…

,

w

i

-

1

)

)

where L f refers to a loss function of the LM model; f refers to the LM model; θ f refers to parameters of the LM; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the LM f with a given utterance U={w 0 ,w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;

wherein the LM of the chatbot is fine-tuned by:

sending, by a client terminal of a user, a first query data to the chatbot;

obtaining, by the fake attacker model, a first content of the first query data;

generating, by the IM of the chatbot, a response data having a response content corresponding to the content of the query data;

obtaining by the fake attacker model, the response content;

sending by the client terminal, a second query data to the chatbot;

obtaining, by the fake attacker model, a second content of the second query data; and

outputting by the fake attacker model, a prediction data having a predicted persona of the user according to the first content the response content and the second content.

8 . The server of claim 7 , wherein the defense objectives comprising one or a combination of following:

a KL (Kullback-Leibler) loss; and

a MI (Mutual Information) loss.

9 . The server of claim 8 , wherein an objective function of the KL loss is presented by a formula below:

L

kl

(

u

;

,

θ

f

)

=

-

1

C

⁢

∑

k

=

0

C

-

1

Pr

⁡

(

k

⁢

❘

"\[LeftBracketingBar]"

f

⁡

(

u

)

,

)

where L kl refers to a loss function of the KL loss; refers to parameters of the fake attacker; u refers to utterance; k refers to personal attribute label index; C refers to a total number of predefined personal attributes; and f(u) refers to hidden states of the chatbot.

10 . The server of claim 8 , wherein an objective function of the MI loss is presented by a formula below:

min

θ

f

max

Ψ

⁢

𝔼

q

⁡

(

f

⁡

(

u

)

)

[

log

⁢

p

Ψ

(

s

⁢

❘

"\[LeftBracketingBar]"

f

⁡

(

u

)

)

]

where E q refers to a loss function of the KL loss; p Ψ (s|f(u)) refers to a distribution function used to approximate q(s|f(u)) which refers to probability distribution of model f parameterized by θ f ; Ψ refers to the attacker model that manages to infer s from f(u).

11 . The server of claim 7 , wherein the fake attacker comprises:

a projection layer comprising a plurality of fully connected layers; and

a softmax activation function layer.

12 . The server of claim 11 , wherein a loss function of the fake attacker is presented by a formula below:

( u kj ,s kj ; )= CE ( p ( f ( u kj ), s kj )

where is the loss function of the fake attacker; u kj represents j-th utterance from k-th conversation; CE refers to cross-entropy between personal attribute label s kj and personal attribute predictor's output p (f(u kj )); and represents parameters of a persona predictor model p of the fake attacker.

13 . A computer-implemented method for preventing leaking a personal privacy from a service under personal privacy inference attack, wherein the service is provided via a neural network executed by a processor of a server, the method comprises:

training, by the processor, a main algorithm model of the service according to utility objectives;

applying, by the processor, one or more defense objectives with attribute predictor to fine-tune the main algorithm model of the service by using a fake attacker model and pre-defined attributes with annotated datasets, wherein the fake attacker is used as a rehearsal to update the main algorithm model in place of the real attacker model details for updating the main algorithm model during actual attacks; and

using, by the processor, the ta fine-tuned main algorithm model on the service to defend inference attack, such that the personal privacy of content inputted and sent to the service cannot be predicted by an external predictor; and a security level of the service is assured;

wherein the utility objectives comprising a LM loss; and

wherein the LM loss is an objective function of the main algorithm model, and the objective function is presented by a formula below:

L

f

(

U

;

θ

f

)

=

-

∑

i

=

1

|

U

|

log

⁡

(

P

⁢

r

⁡

(

w

i

|

c

,

w

0

,

w

1

,

…

,

w

i

-

1

)

)

where L f refers to a loss function of the main algorithm model; f refers to the main algorithm model; θ f refers to parameters of the main algorithm model; w i refers to i th word of a sentence; Pr(w i |c, w 0 , w 1 , . . . , w i-1 ) refers to a probability distribution for the main algorithm model f with a given utterance U={w 0 , w 1 , . . . , w |U|-1 }; c refers to a previous content in a private conversations D;

wherein the main algorithm model of the chatbot is fine tuned by:

sending, by a client terminal of a user, a first query data to the chatbot;

obtaining, by the fake attacker model, a first content of the first query data;

generating, by the main algorithm model of the chatbot, a response data having a response content corresponding to the content of the query data;

obtaining by the fake attacker model, the response content;

sending, by the client terminal, a second query data to the chatbot;

obtaining, by the fake attacker model a second content of the second query data, and

outputting by the fake attacker model, a prediction data having a predicted persona of the user according to the first content, the response content and the second content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: SONG, YANGQIU; LI, HAORAN
To: THE HONG KONG UNIVERSITY OF SCIENCE AND TECHNOLOGY
Reel/Frame 061729/0496 →
Continuity (2)
Provisional Application 63278503 · Nov 12, 2021
Related Publication 20230153460A1 · May 18, 2023
References Cited (22)
US 11132453B2 · Wang · 2021 [cited by examiner]
US 20210064760A1 · Sharma · 2021 [cited by examiner]
Congzheng Song, Ananth Raghunathan “Information Leakage in Embedding Models” [retrieved on Aug. 27, 2025] Retrieved from the Internet <URL: https://arxiv.org/abs/2004.00053> (Year: 2020). [cited by examiner]
Rahman, M. A., Rahman, T., Laganière, R., Mohammed, N., & Wang, Y. (2018). Membership inference attack against differentially private deep learning model. Trans. Data Priv., 11(1), 61-79. (Year: 2018). [cited by examiner]
Maximin Coavoux, Shashi Narayan, Shay B. Cohen “Privacy-preserving Neural Representations of Text” [retrieved on Aug. 27, 2025] Retrieved from the Internet <URL: https://arxiv.org/abs/1808.09408> (Year: 2021). [cited by examiner]
M. Rao et al., “Do as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Can… [cited by examiner]
Yizhe Zhang et al., “DIALOGPT: Large-Scale Generative Pre-training for Conversational Response Generation,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 270-278. [cited by applicant]
Thomas Wolf et al., “TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents,” Association for the Advancement of Artificial Intelligence, 2019. [cited by applicant]
Weizhou Shen et al., “DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion Recognition,” The Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021, pp. 13789-13797. [cited by applicant]
Alec Radford et al., “Language Models are Unsupervised Multitask Learners,” 2019. [cited by applicant]
Zhilin Yang et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” 33rd Conference on Neural Information Processing Systems, 2019. [cited by applicant]
Nicholas Carlini et al., “Extracting Training Data from Large Language Models,” USENIX Security Symposium, 2021. [cited by applicant]
Cynthia Dwork et al., “The Algorithmic Foundations of Differential Privacy,” Foundations and Trends in Theoretical Computer Science, 2014, vol. 9. [cited by applicant]
Sean Welleck et al., “Neural Text Generation With Unlikelihood Training,” 2019. [cited by applicant]
Congzheng Song et al., “Overlearning Reveals Sensitive Attributes,” 8th International Conference on Learning Representations, 2020. [cited by applicant]
Anna Tigunova et al., “Listening between the Lines: Learning Personal Attributes from Conversations,” International World Wide Web Conference, 2019. [cited by applicant]
Chien-Sheng Wu et al., “Getting To Know You: User Attribute Extraction from Dialogues,” LREC, 2020. [cited by applicant]
Eric Lehman et al., “Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?” Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn… [cited by applicant]
Santiago Zanella-Béguelin et al., “Analyzing Information Leakage of Updates to Natural Language Models,” Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, 2020, pp. 363-375. [cited by applicant]
Mohammad Malekzadeh et al., “Honest-but-Curious Nets: Sensitive Attributes of Private Inputs Can Be Secretly Coded into the Classifiers' Outputs,” ACM Conference on Computer and Communications Security, 2021. [cited by applicant]
Yukun Zhu et al., “Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books,” 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 19-27. [cited by applicant]
Xudong Pan et al., “Privacy Risks of General-Purpose Language Models, ” 2020 IEEE Symposium on Security and Privacy, 2020, pp. 1314-1331. [cited by applicant]