IP Library Granted Patent US 11,681,817
Granted Patent B2
US 11,681,817 · App. 16/582,318 · Granted Jun 20, 2023

System and method for implementing attribute classification for PII data

Inventors: Vijaya Kadiyala (Hyderabad, IN); Anil Kumar Gannamani (Hyderabad, IN); Swarna Bhagath Irukulla (Hyderabad, IN)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F21/6209G06F21/6245G06F40/10G06F40/205G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,681,817
App. No.
16/582,318
Filed
Sep 25, 2019
Granted
Jun 20, 2023
Kind
B2
Art Unit
2177
USPC
726/26
Abstract

An embodiment of the present invention is directed to classifying attributes into respective PI/PG categories based on metadata. An embodiment of the present invention may classify each attribute into PII/Non-PII and then into various Protection group codes that define access, roles permissions, privileges and/or other action. An embodiment of the present invention may leverage various statistical techniques, natural language processing (NLP) methods and different combinations of algorithms customized to improve prediction accuracies of a classifier model.

Claims (44)

1. A system that performs data classification for personally identifiable information (PII) data, the system comprising:

a database system that stores attribute data and corresponding metadata;

an interactive user interface configured to receive user input via a communication network; and

a computer processor, coupled to the database system and the communication network, configured to:

create machine learning algorithms;

define at least one of the machine learning algorithms by a hyperplane that distinctly classifies one or more data points;

receive data relating to one or more attributes of the attribute data;

apply tokenization to the data relating to the one or more attributes of the attribute data, to identify a sequence of words;

apply, to the data relating to the one or more attributes of the attribute data, a weighting technique that represents an importance associated with a word;

identify corresponding metadata associated with the data relating to the one or more attributes of the attribute data, wherein the identified corresponding metadata includes values determined by the applying the weighting technique; and

classify the data relating to the one or more attributes of the attribute data into non-PII data and the PII data based on the identified corresponding metadata and further classifying the PII data into one of a plurality of protection groups, each protection group identifying access permissions,

wherein the classify is based on statistical techniques, the machine learning algorithms, and natural language processing.

2. The system of claim 1 , wherein the computer processor is configured to:

apply a stemming process to the data relating to the one or more attributes of the attribute data, to reduce inflected words to a stem form.

3. The system of claim 1 , wherein the computer processor is configured to:

apply a grouping of words as a single item.

4. The system of claim 3 , wherein the grouping comprises a lemmatization process.

5. The system of claim 1 , wherein the weighting technique comprises term frequency-inverse document frequency (TF-IDF) weight.

6. The system of claim 5 , wherein the term frequency-inverse document frequency (TF-IDF) weight comprises a first term that measures how frequently a term appears in a document.

7. The system of claim 6 , wherein the term frequency-inverse document frequency (TF-IDF) weight comprises a second term that measures how important a term is.

8. The system of claim 1 , wherein the computer processor is configured to: apply a Synthetic Minority Oversampling Technique to increase a dataset in a balanced manner.

9. The system of claim 1 , wherein the computer processor is further configured to: plot the data relating to the one or more attributes of the attribute data as the one or more data points in the hyperplane, wherein each feature of an attribute of the one or more attributes of the attribute data has a one-to-one correspondence with each dimension in the hyperplane.

10. The system of claim 1 , wherein the interactive user interface is a microservice that is further configured to: utilize REST API to consume the user input by providing, to the computer processor, the user input as the data relating to the one or more attributes of the attribute data.

11. A method that performs data classification for personally identifiable information (PII) data, the method comprising:

storing attribute data and corresponding metadata

creating machine learning algorithms;

defining at least one of the machine learning algorithms by a hyperplane that distinctly classifies one or more data points;

receiving data relating to one or more attributes of the attribute data;

applying tokenization to the data relating to the one or more attributes of the attribute data, to identify a sequence of words;

applying, to the data relating to the one or more attributes of the attribute data, a weighting technique that represents an importance associated with a word;

identifying corresponding metadata associated with the data relating to the one or more attributes of the attribute data, wherein the identified corresponding metadata includes values determined by the applying the weighting technique; and

classifying the data relating to the one or more attributes of the attribute data into non-PII data and the PII data based on the identified corresponding metadata and further classifying the PII data into one of a plurality of protection groups, each protection group identifying access permissions,

wherein the classifying is based on statistical techniques, the machine learning algorithms, and natural language processing.

12. The method of claim 11 , further comprising:

applying a stemming process to the data relating to the one or more attributes of the attribute data, to reduce inflected words to a stem form.

13. The method of claim 11 , further comprising:

applying a grouping of words as a single item.

14. The method of claim 13 , wherein the grouping comprises a lemmatization process.

15. The method of claim 11 , wherein the weighting technique comprises term frequency-inverse document frequency (TF-IDF) weight.

16. The method of claim 15 , wherein the term frequency-inverse document frequency (TF-IDF) weight comprises a first term that measures how frequently a term appears in a document.

17. The method of claim 16 , wherein the term frequency-inverse document frequency (TF-IDF) weight comprises a second term that measures how important a term is.

18. The method of claim 11 , further comprising: applying a Synthetic Minority Oversampling Technique to increase a dataset in a balanced manner.

19. The method of claim 11 , further comprising: plotting the data relating to the one or more attributes of the attribute data as the one or more data points in the hyperplane, wherein each feature of an attribute of the one or more attributes of the attribute data has a one-to-one correspondence with each dimension in the hyperplane.

20. The method of claim 11 , further comprising: utilizing, by an interactive user interface that is a microservice, REST API to consume the user input by providing, to a computer processor, the user input as the data relating to the one or more attributes of the attribute data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: KADIYALA, VIJAYA; GANNAMANI, ANIL KUMAR; IRUKULLA, SWARNA BHAGATH
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 053631/0180 →
Continuity (1)
Related Publication 20210089667A1 · Mar 25, 2021
Cited By (1)
US 12,373,530