IP Library Granted Patent US 12,706,198
Granted Patent B2
US 12,706,198 · App. 18/828,410 · Granted Aug 11, 2026

Privacy protection tuning for LLMs in medical decision making

Inventors: Wei Cheng (Princeton Junction, NJ); Wenchao Yu (Plainsboro, NJ); Yanchi Liu (Monmouth Junction, NJ); Xujiang Zhao (Hillsborough, NJ); Haifeng Chen (West Windsor, NJ); Yijia Xiao (Los Angeles, CA)
Assignee: NEC Corporation
G16H20/00G06F21/6245G06F40/284G16H10/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,706,198
App. No.
18/828,410
Filed
Sep 9, 2024
Granted
Aug 11, 2026
Kind
B2
Art Unit
2433
USPC
726/26
Abstract

Methods and systems include annotating a set of training data to indicate tokens that are sensitive. Instructions are generated based on the training data, including original token sequences and respective substituted token sequences. A language model is fine-tuned using the instructions with a penalty-based loss function to generate a privacy-protected language model.

Claims (244)

1 . A computer-implemented method, comprising:

annotating a set of training data to indicate tokens that are sensitive;

generating instructions based on the training data, including original token sequences and respective substituted token sequences; and

fine-tuning a language model using the instructions with a penalty-based loss function to generate a privacy-protected language model, wherein the penalty-based loss includes separate unigram and bigram terms, wherein the unigram term is:

l

1

gram

(

s

,

k

)

=

w

1

PII

Θ

1

P

(

w

1

PII

"\[LeftBracketingBar]"

{

w

i

}

i

=

1

k

-

1

)

and where the bigram term is

l

2

gram

(

s

,

k

)

=

(

w

1

PII

,

w

2

PII

)

Θ

2

P

(

w

1

PII

"\[LeftBracketingBar]"

{

w

i

}

i

=

1

k

-

1

)

P

(

w

2

PII

"\[LeftBracketingBar]"

{

w

i

}

i

=

1

k

)

where s is a sequence, k is a position in the sequence

w

1

PII

and

w

2

PII

are tokens associated with sensitive information, w 1 is a term of the sequence, Θ 1 is a set of unigrams associated with sensitive information, Θ 2 is a set of bigrams associated with sensitive information, and P is a probability function.

2 . The method of claim 1 , wherein annotating the set of training data includes generating a privacy label sequence for each original sequence in the training data with a binary sensitivity indicator for each token of respective original sequences.

3 . The method of claim 1 , wherein the substituted token sequences include the tokens of the original token sequences but with sensitive tokens being replaced by a placeholder.

4 . The method of claim 1 , wherein generating the instructions includes generating a positive example and a negative example based on an original token sequence and a respective substituted token sequence.

5 . The method of claim 1 , wherein the sensitive tokens relate to personally identifiable information.

6 . The method of claim 1 , further comprising providing a patient's medical records to the privacy-protected language model to aid in medical decision making, wherein the training data includes medical records.

7 . The method of claim 1 , wherein the language model is a pre-trained machine learning model that processes natural language inputs.

8 . The method of claim 6 , further comprising automatically altering the patient's treatment responsive to an output of the privacy-protected language model.

9 . A system, comprising:

a hardware processor; and

a memory that stores a computer program that, when executed by the hardware processor, causes the hardware processor to:

annotate a set of training data to indicate tokens that are sensitive;

generate instructions based on the training data, including original token sequences and respective substituted token sequences; and

fine-tune a language model using the instructions with a penalty-based loss function to generate a privacy-protected language model, wherein the penalty-based loss includes separate unigram and bigram terms, wherein the unigram term is:

l

1

gram

(

s

,

k

)

=

w

1

PII

Θ

1

P

(

w

1

PII

|

{

w

i

}

i

=

1

k

-

1

)

and where the bigram term is

l

2

gram

(

s

,

k

)

=

(

w

1

PII

,

w

2

PII

)

Θ

2

P

(

w

1

PII

"\[LeftBracketingBar]"

{

w

i

}

i

=

1

k

-

1

)

P

(

w

2

PII

"\[LeftBracketingBar]"

{

w

i

}

i

=

1

k

)

where s is a sequence, k is a position in the sequence,

w

1

PII

and

w

2

PII

 are tokens associated with sensitive information, w 1 is a term of the sequence, Θ 1 is a set of unigrams associated with sensitive information, Θ 2 is a set of bigrams associated with sensitive information, and P is a probability function.

10 . The system of claim 9 , wherein the computer program further causes the hardware processor to generate a privacy label sequence for each original sequence in the training data with a binary sensitivity indicator for each token of respective original sequences.

11 . The system of claim 9 , wherein the substituted token sequences include the tokens of the original token sequences but with sensitive tokens being replaced by a placeholder.

12 . The system of claim 9 , wherein the computer program further causes the hardware processor to generate a positive example and a negative example based on an original token sequence and a respective substituted token sequence.

13 . The system of claim 9 , wherein the sensitive tokens relate to personally identifiable information.

14 . The system of claim 9 , wherein the computer program further causes the hardware processor to provide a patient's medical records to the privacy-protected language model to aid in medical decision making, wherein the training data includes medical records.

15 . The system of claim 9 , wherein the language model is a pre-trained machine learning model that processes natural language inputs.

16 . The system of claim 14 , wherein the computer program further causes the hardware processor to automatically alter the patient's treatment responsive to an output of the privacy-protected language model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2026
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 075069/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2024
From: CHENG, WEI; YU, WENCHAO; LIU, YANCHI; ZHAO, XUJIANG; CHEN, HAIFENG; XIAO, YIJIA
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 068528/0468 →
Continuity (2)
Provisional Application 63539623 · Sep 21, 2023
Related Publication 20250104824A1 · Mar 27, 2025
References Cited (38)
US 12182179B1 · Aravamudan · 2024 [cited by examiner]
US 12259864B1 · Aravamudan · 2025 [cited by examiner]
US 20200349271A1 · Binkley · 2020 [cited by examiner]
US 20230259787A1 · David · 2023 [cited by examiner]
US 20240363247A1 · Attia · 2024 [cited by examiner]
WO WO2025145165A1 · 2025 [cited by examiner]
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., . . . & Kaplan, J. (Apr. 12, 2022). Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:22… [cited by applicant]
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., . . . & Fung, P. (Feb. 8, 2023). A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint… [cited by applicant]
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., . . . & Raffel, C. (Aug. 11, 2021). Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security… [cited by applicant]
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., & Zhang, C. (Feb. 15, 2022). Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646. [cited by applicant]
Chiang, W. L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., . . . & Xing, E. P. (Mar. 30, 2023). Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. Imsys. org, 2(3), 6. [cited by applicant]
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., & Tang, J. (Mar. 17, 2021). Glm: General language model pretraining with autoregressive blank infilling. arXiv preprint arXiv:2103.10360. [cited by applicant]
Gehman, S., Gururangan, S., Sap, M., Choi, Y., & Smith, N. A. (Sep. 24, 2020). Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462. [cited by applicant]
Hendrycks, D., Mazeika, M., Kadavath, S., & Song, D. (Dec. 8, 2019). Using self-supervised learning can improve model robustness and uncertainty. Advances in neural information processing systems, 32. [cited by applicant]
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., . . . & Sifre, L. (Mar. 29, 2022). Training compute-optimal large language models. arXiv preprint arXiv:2203.15556. [cited by applicant]
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., . . . & Sifre, L. (Dec. 6, 2022). An empirical analysis of compute-optimal large language model training. Advances in Neural Information … [cited by applicant]
Devlin, J., Chang, M., Lee, K., Toutanova, K. (May 24, 2019). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805v2. [cited by applicant]
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., . . . & Perez, E. (Jul. 3, 2023). Pretraining language models with human preferences. In International Conference on Machine Learning (pp. 17506-17… [cited by applicant]
Lin, S., Hilton, J., & Evans, O. (Sep. 8, 2021). Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958. [cited by applicant]
Meng, Y., Michalski, M., Huang, J., Zhang, Y., Abdelzaher, T., & Han, J. (Jul. 3, 2023). Tuning language models as training data generators for augmentation-enhanced few-shot learning. In International Conference on Mac… [cited by applicant]
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., . . . & McAleese, N. (Mar. 21, 2022). Teaching language models to support answers with verified quotes. arXiv preprint arXiv:2203.11147. [cited by applicant]
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., . . . & Lowe, R. (Dec. 6, 2022). Training language models to follow instructions with human feedback. Advances in neural information processing sy… [cited by applicant]
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., . . . & Zhu, R. J. (May 22, 2023). Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048. [cited by applicant]
Ramasesh, V. V., Lewkowycz, A., & Dyer, E. (Apr. 25, 2022). Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations. [cited by applicant]
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., . . . & Natarajan, V. (Jul. 12, 2023). Large language models encode clinical knowledge. Nature, 620(7972), 172-180. [cited by applicant]
Solaiman, I., & Dennison, C. (Dec. 6, 2021). Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34, 5861-5873. [cited by applicant]
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., . . . & Hashimoto, T. B. (Mar. 13, 2023). Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca. [cited by applicant]
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., . . . & Lample, G. (Feb. 27, 2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971. [cited by applicant]
Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., & Ho, A. (Oct. 26, 2022). Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325. [cited by applicant]
Vu, T., Barua, A., Lester, B., Cer, D., Iyyer, M., & Constant, N. (May 25, 2022). Overcoming catastrophic forgetting in zero-shot cross-lingual generation. arXiv preprint arXiv:2205.12647. [cited by applicant]
Wang, Y., Zhong, W., Li, L., Mi, F., Zeng, X., Huang, W., . . . & Liu, Q. (Jul. 24, 2023). Aligning large language models with human: A survey. arXiv preprint arXiv:2307.12966. [cited by applicant]
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., . . . & Huang, P. S. (Sep. 15, 2021). Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445. [cited by applicant]
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., . . . & Mann, G. (Mar. 30, 2023). Bloomberggpt: A large language model for finance. arXiv preprint arXiv:2303.17564. [cited by applicant]
Xu, J., Ju, D., Li, M., Boureau, Y. L., Weston, J., & Dinan, E. (Oct. 14, 2020). Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079. [cited by applicant]
Yu, W., Pang, T., Liu, Q., Du, C., Kang, B., Huang, Y., . . . & Yan, S. (Jul. 3, 2023). Bag of tricks for training data extraction from language models. In International Conference on Machine Learning (pp. 40306-40320).… [cited by applicant]
Zhang, X., Li, S., Yang, X., Tian, C., Qin, Y., & Petzold, L. R. (May 22, 2023). Enhancing small medical learners with privacy-preserving contextual prompting. arXiv preprint arXiv:2305.12723. [cited by applicant]
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., . . . & Irving, G. (Jan. 8, 2020). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593. [cited by applicant]
Ziegler, D., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., . . . & Thomas, N. (Dec. 6, 2022). Adversarial training for high-stakes reliability. Advances in Neural Information Processing Systems, 35, 9274-… [cited by applicant]