IP Library Granted Patent US 12,293,003
Granted Patent B2
US 12,293,003 · App. 18/654,684 · Granted May 6, 2025

Machine learning modeling to identify sensitive data

Inventors: Shubhanshu Gupta (Singapore, SG); Ashish Awasthi (Singapore, SG); Amaruvi Devanathan (Chennai, IN); Mallapu Raghavulu Surya Prakash (Singapore, SG)
Assignee: CITIBANK, N.A.
G06F21/6254G06F16/221G06F16/3347G06F16/335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,003
App. No.
18/654,684
Granted
May 6, 2025
Kind
B2
Abstract

Methods and systems herein identify and redact personally identifiable information. A PII sensitivity detection framework includes multiple layers where each layer corresponds to a computer model. The framework analyzes data stored within different data tables and predicts whether a data column includes PII. The first layer corresponds to an artificial intelligence model that analyzes each column metadata and predicts a first score indicative of a likelihood of PII. The second layer corresponds to a rule-based computer model that uses various rules to determine a second score indicative of a likelihood of PII for each column. The third layer corresponds to a column content model that analyzes content of each column using various natural language processing techniques to generate a third score indicative of a likelihood of PII. The framework masks data being presented to a user based on the scores generated via execution of one or more of the layers.

Claims (36)

1. A method comprising:

executing, by a processor, a first artificial intelligence model to generate a first score corresponding a first likelihood of a set of text including personally identifiable information;

executing, by the processor, a second artificial intelligence model to generate a second score corresponding to a second likelihood of the set of text including personally identifiable information, the second artificial intelligence model determining the second score based on a cardinality value or a length value associated with the set of text; and

masking, by the processor, at least a portion of the set of text likely to include personally identifiable information in accordance with the first score and the second score.

2. The method of claim 1 , wherein the first artificial intelligence model uses metadata associated with the set of text to determine the first likelihood.

3. The method of claim 1 , wherein masking at least the portion of the set of text corresponds to:

instructing, by the processor, a webserver to redact at least the portion of the set of text.

4. The method of claim 1 , wherein the first artificial intelligence model is configured to ingest a vector associated with the set of text to predict the first score, wherein the vector is generated using at least one of one-hot encoding, count vectorization, or term frequency method.

5. The method of claim 1 , further comprising:

executing, by the processor, a third artificial intelligence model to generate a fourth score corresponding to a fourth likelihood of the set of text including personally identifiable information, the third artificial intelligence model determining the fourth score based on executing a natural language processing protocol.

6. The method of claim 1 , wherein the first artificial intelligence model uses metadata associated with the set of text to generate the first score, wherein the metadata corresponds to at least one of a geographical region, a classification, a data type, a defined number of characters for at least one data record, or a description associated with the set of text.

7. The method of claim 1 , wherein the processor masks at least the portion of the set of text when a user viewing at least the portion of the text has a user attribute that satisfies a threshold.

8. The method of claim 1 , wherein masking at least the portion of the set of text corresponds to revising, by the processor, a data record that corresponds to at least the portion of the set of text.

9. A system comprising:

a server comprising a processor and a non-transitory computer-readable medium containing instructions that when executed by the processor causes the processor to perform operations comprising:

execute a first artificial intelligence model to generate a first score corresponding a first likelihood of a set of text including personally identifiable information;

execute a second artificial intelligence model to generate a second score corresponding to a second likelihood of the set of text including personally identifiable information, the second artificial intelligence model determining the second score based on a cardinality value or a length value associated with the set of text;

mask at least a portion of the set of text likely to include personally identifiable information in accordance with the first score and the second score.

10. The system of claim 9 , wherein the first artificial intelligence model uses metadata associated with the set of text to determine the first likelihood.

11. The system of claim 9 , wherein masking at least the portion of the set of text corresponds to instructing a webserver to redact at least the portion of the set of text.

12. The system of claim 9 , wherein the first artificial intelligence model is configured to ingest a vector associated with the set of text to predict the first score, wherein the vector is generated using at least one of one-hot encoding, count vectorization, or term frequency method.

13. The system of claim 9 , wherein the instructions further cause the processor to:

execute a third artificial intelligence model to generate a fourth score corresponding to a fourth likelihood of the set of text including personally identifiable information, the third artificial intelligence model determining the fourth score based on executing a natural language processing protocol.

14. The system of claim 9 , wherein the first artificial intelligence model uses metadata associated with the set of text to generate the first score, wherein the metadata corresponds to at least one of a geographical region, a classification, a data type, a defined number of characters for at least one data record, or a description associated with the set of text.

15. The system of claim 9 , wherein the processor masks at least the portion of the set of text when a user viewing at least the portion of the text has a user attribute that satisfies a threshold.

16. The system of claim 9 , wherein masking at least the portion of the set of text corresponds to revising, by the processor, a data record that corresponds to at least the portion of the set of text.

17. A system comprising:

a first artificial intelligence model;

a second artificial intelligence model;

a server having at least one processor in communication with the first artificial intelligence model and the second artificial intelligence model, the server configured to:

execute the first artificial intelligence model to generate a first score corresponding a first likelihood of a set of text including personally identifiable information;

execute the second artificial intelligence model to generate a second score corresponding to a second likelihood of the set of text including personally identifiable information, the second artificial intelligence model determining the second score based on a cardinality value or a length value associated with the set of text;

mask at least a portion of the set of text likely to include personally identifiable information in accordance with the first score and the second score.

18. The system of claim 17 , wherein the first artificial intelligence model uses metadata associated with the set of text to determine the first likelihood.

19. The system of claim 17 , wherein masking at least the portion of the set of text corresponds to instructing a webserver to redact at least the portion of the set of text.

20. The system of claim 17 , wherein the first artificial intelligence model is configured to ingest a vector associated with the set of text to predict the first score, wherein the vector is generated using at least one of one-hot encoding, count vectorization, or term frequency method.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2024
From: GUPTA, SHUBHANSHU; AWASTHI, ASHISH; DEVANATHAN, AMARUVI; PRAKASH, MALLAPU RAGHAVULU SURYA
To: CITIBANK, N.A.
Reel/Frame 067311/0358 →
Continuity (2)
Continuation 17476388 · Sep 15, 2021
Related Publication 20240289492A1 · Aug 29, 2024
References Cited (9)
US 11630853B2 · Hawco · 2023 [cited by examiner]
US 11755766B2 · Kulkarni · 2023 [cited by examiner]
US 20180165475A1 · Veeramachaneni · 2018 [cited by examiner]
US 20180285599A1 · Praveen · 2018 [cited by examiner]
US 20210067542A1 · Linder · 2021 [cited by examiner]
US 20210073412A1 · Kvochko · 2021 [cited by examiner]
US 20210125089A1 · Nickl · 2021 [cited by examiner]
US 20220043935A1 · Brannon · 2022 [cited by examiner]
US 20220198044A1 · Madhavan · 2022 [cited by examiner]