IP Library › Granted Patent US 12,591,708
Granted Patent B2
US 12,591,708 · App. 17/095,185 · Granted Mar 31, 2026

Personal data anonymization with model verification

Inventors: Aaron Beach (Lakewood, CO); Liyuan Zhang (San Mateo, CA); Tiffany Callahan (Boulder, CO); Mengda Liu (Pittsburgh, PA); Shruti Vempati (Denver, CO)
Assignee: Twilio Inc.
G06F21/6254G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,708
App. No.
17/095,185
Granted
Mar 31, 2026
Kind
B2
Abstract

Anonymization is the process to remove personal information from the data. Once the data is anonymized, the data may be used for creating machine-learning models without the risk of invading anyone's privacy. One of the keys to data anonymization is to make sure that the data is really anonymized so nobody could use the anonymized data to obtain private information. Different techniques for data anonymization are presented. Further, the data anonymization techniques are tested for true anonymization by comparing the results from these techniques to a random method of guessing. If the difference in the results is below a predetermined threshold margin, the data anonymization techniques are safe and ready for use.

Claims (46)

1 . A computer-implemented method comprising:

generating training data associated with electronic messages to be transmitted by a plurality of users, wherein the training data correlates anonymized communications data created by anonymizing the electronic messages to be transmitted by the plurality of users to personally identifiable information (PII) extracted from the electronic messages to be transmitted by the plurality of users, wherein the electronic messages comprise one or more of email messages or short message service (SMS) messages, and wherein anonymizing the electronic messages comprises redacting text of a communication to delete one or more PII values and embedding the redacted text into a vector;

training, by one or more processors and based on the training data associated with the electronic messages to be transmitted by the plurality of users, a machine-learning program to generate a PII detection model that extracts PII values from input electronic messages;

extracting, by the PII detection model executing on the one or more processors, PII values from anonymized testing electronic messages;

determining differences between the extracted PII values obtained from the PII detection model and randomly generated PII values obtained from a random value model that randomly generates PII values using a list of PII values found in the anonymized testing electronic messages based on frequencies of appearance; and

responsive to the determined differences between the extracted PII values and the randomly generated PII values being below a predetermined confidence level, determining that anonymization of the electronic messages is valid based on the PII detection model; and

causing spam detection to be performed using the electronic messages anonymized based on the PII detection model.

2 . The method as recited in claim 1 , wherein the training data comprises a matrix, each row of the matrix including PII extracted from a corresponding communication and including a vector that anonymizes the corresponding communication.

3 . The method as recited in claim 1 , wherein the extracted PII values include a PII value extracted from a text of an email message.

4 . The method as recited in claim 1 , wherein the random value model is trained to randomly generate a PII value based on the list of PII values found in the anonymized testing electronic messages based on the frequencies of appearance.

5 . The method as recited in claim 1 , wherein the determining of the differences between the extracted PII values and the randomly generated PII values is based on a difference between a first accuracy in first standard deviations of the PII detection model and a second accuracy in second standard deviations of the random value model.

6 . The method as recited in claim 1 , further comprising:

responsive to the determined differences between the extracted PII values and the randomly generated PII values transgressing the predetermined confidence level, determining that the anonymized testing electronic messages were insufficiently anonymized based on the PII detection model, wherein the determining that the anonymized testing electronic messages were insufficiently anonymized comprises:

determining a difference between a first performance of the PII detection model and a second performance of the random value model; and

determining that the first and second performances are equivalent when the difference is below a predetermined level.

7 . The method as recited in claim 1 , wherein the extracted PII values include a first name and a last name.

8 . A system comprising:

a memory comprising instructions; and

one or more computer processors, wherein the instructions, when executed by the one or more computer processors, cause the system to perform operations comprising:

generating training data associated with electronic messages to be transmitted by a plurality of users, wherein the training data correlates anonymized communications data created by anonymizing the electronic messages to be transmitted by the plurality of users to personally identifiable information (PII) extracted from the electronic messages to be transmitted by the plurality of users, wherein the electronic messages comprise one or more of email messages or short message service (SMS) messages, and wherein anonymizing the electronic messages comprises redacting text of a communication to delete one or more PII values and embedding the redacted text into a vector;

training, based on the training data associated with the electronic messages to be transmitted by the plurality of users, a machine-learning program to generate a PII detection model that extracts PII values from input electronic messages;

extracting, by the PII detection model executing on the one or more computer processors, PII values from anonymized testing electronic messages;

determining differences between the extracted PII values obtained from the PII detection model and randomly generated PII values obtained from a random value model that randomly generates PII values using a list of PII values found in the anonymized testing electronic messages based on frequencies of appearance; and

responsive to the determined differences between the extracted PII values and the randomly generated PII values being below a predetermined confidence level, determining that anonymization of the electronic messages is valid based on the PII detection model; and

causing spam detection to be performed using the electronic messages anonymized based on the PII detection model.

9 . The system as recited in claim 8 , wherein the training data comprises a matrix, each row of the matrix including PII extracted from a corresponding communication and including a vector that anonymizes the corresponding communication.

10 . The system as recited in claim 8 , wherein the random value model is trained to randomly generate a PII value based on the list of PII values found in the anonymized testing electronic messages based on the frequencies of appearance.

11 . The system as recited in claim 8 , wherein the determining of the differences between the extracted PII values and the randomly generated PII values is based on a difference between a first accuracy in first standard deviations of the PII detection model and a second accuracy in second standard deviations of the random value model.

12 . A non-transitory machine-readable storage medium including instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

generating training data associated with electronic messages to be transmitted by a plurality of users, wherein the training data correlates anonymized communications data created by anonymizing the electronic messages to be transmitted by the plurality of users to personally identifiable information (PII) extracted from the electronic messages to be transmitted by the plurality of users, wherein the electronic messages comprise one or more of email messages or short message service (SMS) messages, and wherein anonymizing the electronic messages comprises redacting text of a communication to delete one or more PII values and embedding the redacted text into a vector;

training, based on the training data associated with the electronic messages to be transmitted by the plurality of users, a machine-learning program to generate a PII detection model that extracts PII values from input electronic messages;

extracting, by the PII detection model executing on the one or more processors, PII values from anonymized testing electronic messages;

determining differences between the extracted PII values obtained from the PII detection model and randomly generated PII values obtained from a random value model that randomly generates PII values using a list of PII values found in the anonymized testing electronic messages based on frequencies of appearance; and

responsive to the determined differences between the extracted PII values and the randomly generated PII values being below a predetermined confidence level, determining that anonymization of the electronic messages is valid based on the PII detection model; and

causing spam detection to be performed using the electronic messages anonymized based on the PII detection model.

13 . The non-transitory machine-readable storage medium as recited in claim 12 , wherein the training data comprises a matrix, each row of the matrix including PII extracted from a corresponding communication and including a vector that anonymizes the corresponding communication.

14 . The non-transitory machine-readable storage medium as recited in claim 12 , wherein the random value model is trained to randomly generate a PII value based on the list of PII values found in the anonymized testing electronic messages based on the frequencies of appearance.

15 . The non-transitory machine-readable storage medium as recited in claim 12 , wherein the determining of the differences between the extracted PII values and the randomly generated PII values is based on a difference between a first accuracy in first standard deviations of the PII detection model and a second accuracy in second standard deviations of the random value model.

16 . The system as recited in claim 8 , the operations further comprising:

responsive to the determined differences between the extracted PII values and the randomly generated PII values transgressing the predetermined confidence level, determining that the anonymized testing electronic messages were insufficiently anonymized based on the PII detection model, wherein the determining that the anonymized testing electronic messages were insufficiently anonymized comprises:

determining a difference between a first performance of the PII detection model and a second performance of the random value model; and

determining that the first and second performances are equivalent when the difference is below a predetermined level.

17 . The non-transitory machine-readable storage medium as recited in claim 12 , the operations further comprising:

responsive to the determined differences between the extracted PII values and the randomly generated PII values transgressing the predetermined confidence level, determining that the anonymized testing electronic messages were insufficiently anonymized based on the PII detection model, wherein the determining that the anonymized testing electronic messages were insufficiently anonymized comprises:

determining a difference between a first performance of the PII detection model and a second performance of the random value model; and

determining that the first and second performances are equivalent when the difference is below a predetermined level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2021
From: BEACH, AARON; ZHANG, LIYUAN; CALLAHAN, TIFFANY; LIU, MENGDA; VEMPATI, SHRUTI
To: TWILIO INC.
Reel/Frame 057124/0218 →
Continuity (1)
Related Publication 20220147654A1 · May 12, 2022
References Cited (30)
US 10565398B2 · Huang · 2020 [cited by examiner]
US 10970414B1 · Lesner · 2021 [cited by examiner]
US 11250162B2 · Praveen · 2022 [cited by examiner]
US 11483294B2 · Karabatis · 2022 [cited by examiner]
US 11755766B2 · Kulkarni · 2023 [cited by examiner]
US 11757816B1 · Lin · 2023 [cited by examiner]
US 20070266079A1 · Criddle · 2007 [cited by examiner]
US 20090240637A1 · Drissi · 2009 [cited by examiner]
US 20100036884A1 · Brown · 2010 [cited by examiner]
US 20100162402A1 · Rachlin · 2010 [cited by examiner]
US 20100318489A1 · De Barros · 2010 [cited by examiner]
US 20140040172A1 · Ling · 2014 [cited by examiner]
US 20140115710A1 · Hughes · 2014 [cited by examiner]
US 20140115715A1 · Pasdar · 2014 [cited by examiner]
US 20140280261A1 · Butler · 2014 [cited by examiner]
US 20170024581A1 · Grubel · 2017 [cited by examiner]
US 20170359313A1 · Livneh · 2017 [cited by examiner]
US 20180255010A1 · Goyal · 2018 [cited by examiner]
US 20180314853A1 · Oliner · 2018 [cited by examiner]
US 20190050599A1 · Canard · 2019 [cited by examiner]
US 20200050966A1 · Enuka · 2020 [cited by examiner]
US 20200082290A1 · Pascale · 2020 [cited by examiner]
US 20210075824A1 · Ibrahim · 2021 [cited by examiner]
US 20210173854A1 · Wilshinsky · 2021 [cited by examiner]
US 20220277097A1 · Cabot · 2022 [cited by examiner]
CN 110430224A · 2019 [cited by examiner]
EP 3340561A1 · 2018 [cited by examiner]
EP 3477530A1 · 2019 [cited by examiner]
WO WO2015077542A1 · 2015 [cited by examiner]
Ren et al., “ReCon: Revealing and Controlling PII Leaks in Mobile Network Traffic”, arXiv, pp. 1-18 (Year: 2016). [cited by examiner]