IP Library Granted Patent US 11,899,807
Granted Patent B2
US 11,899,807 · App. 17/462,983 · Granted Feb 13, 2024

Systems and methods for auto discovery of sensitive data in applications or databases using metadata via machine learning techniques

Inventors: Santosh Chikoti (Monroe Township, NJ); Jeffrey Kessler (Mahopac, NY); Ita B Lamont (North Brunswick, NJ); Saurabh Gupta (Secaucus, NJ)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F21/6218G06F16/35G06F16/383G06F40/151
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,899,807
App. No.
17/462,983
Granted
Feb 13, 2024
Kind
B2
Abstract

A method for auto discovery of sensitive data may include: (1) receiving, at data enrichment computer program in a metadata processing pipeline, raw metadata from a plurality of different data sources; (2) enriching, by the data enrichment computer program, the raw metadata; (3) converting, by the data enrichment computer program, the raw metadata and the enhanced raw metadata into a sentence structure; (4) predicting, by a category prediction computer program in the metadata processing pipeline, a predicted category for the sentence structure; (5) identifying, by a sensitive data mapping computer program, a sensitive data category that is mapped to the predicted category based on a policy mapping rule; (6) determining, by the sensitive data mapping computer program, a risk classification rating for the predicted category; and (7) tagging, by the sensitive data mapping computer program, the data source associated with the metadata based on the risk classification rating.

Claims (36)

1. A method for auto discovery of sensitive data in applications or databases using metadata via machine learning techniques, comprising:

receiving, at data enrichment computer program in a metadata processing pipeline, raw metadata from a plurality of different data sources;

enriching, by the data enrichment computer program, the raw metadata;

converting, by the data enrichment computer program, the raw metadata and the enhanced raw metadata into a sentence structure;

predicting, by a category prediction computer program in the metadata processing pipeline, a predicted category for the sentence structure;

identifying, by a sensitive data mapping computer program, a sensitive data category that is mapped to the predicted category based on a policy mapping rule;

determining, by the sensitive data mapping computer program, a risk classification rating for the predicted category; and

tagging, by the sensitive data mapping computer program, the data source associated with the metadata based on the risk classification rating.

2. The method of claim 1 , wherein the data source is tagged with a highly confidential tag, a confidential tag, a private tag, or a public tag.

3. The method of claim 1 , wherein one of the data sources comprises a database.

4. The method of claim 1 , wherein one of the data sources comprises and application.

5. The method of claim 1 , wherein the sensitive data mapping computer program looks up the predicted category in a sensitive data category database.

6. The method of claim 5 , wherein the sensitive data category database is specific to an organization.

7. The method of claim 1 , further comprising:

receiving, from a user interface, a query for a sensitive data classification; and

returning, to the user interface, a result to the query.

8. The method of claim 4 , further comprising:

combining, by the data enrichment computer program, the sentence structure with application data for the application.

9. The method of claim 1 , wherein the category prediction computer program comprises a trained model binary object, wherein the trained model binary object returns the predicted category for the sentence structure.

10. The method of claim 9 , wherein the trained model binary object is trained with historical data using a supervised learning/training process.

11. A system, comprising:

a plurality of data sources comprising raw metadata;

a metadata harvester that harvests the raw metadata from the plurality of data sources;

a sensitive data database that maps categories of metadata to a sensitive data category;

a metadata processing pipeline that receives the raw metadata from the data harvesters, comprising:

a data enrichment computer program that is configured to enrich the raw metadata and covert the raw metadata and the enhanced raw metadata into a sentence structure;

a category prediction computer program that is configured to predict a predicted category for the sentence structure;

a sensitive data mapping computer program that is configured to identify a sensitive data category in the sensitive data database that is mapped to the predicted category based on a policy mapping rule, determine a risk classification rating for the predicted category, and tag the data source associated with the metadata based on the risk classification rating.

12. The system of claim 11 , wherein the data source is tagged with a highly confidential tag, a confidential tag, a private tag, or a public tag.

13. The system of claim 11 , wherein one of the data sources comprises a database.

14. The system of claim 11 , wherein one of the data sources comprises and application.

15. The system of claim 11 , wherein the sensitive data mapping computer program is further configured to look up the predicted category in a sensitive data category database.

16. The system of claim 15 , wherein the sensitive data category database is specific to an organization.

17. The system of claim 11 , wherein the data enrichment computer program is further configured to combine the sentence structure with application data for the application.

18. The system of claim 11 , wherein the category prediction computer program comprises a trained model binary object, wherein the trained model binary object returns the predicted category for the sentence structure.

19. The system of claim 18 wherein the trained model binary object is trained with historical data using a supervised learning/training process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: CHIKOTI, SANTOSH; KESSLER, JEFFREY; LAMONT, ITA B; GUPTA, SAURABH
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 065323/0692 →
Continuity (2)
Provisional Application 63073572 · Sep 2, 2020
Related Publication 20220067185A1 · Mar 3, 2022