IP Library Granted Patent US 12,373,584
Granted Patent B2
US 12,373,584 · App. 18/358,130 · Granted Jul 29, 2025

Systems and methods for securing data based on discovered relationships

Inventors: Vijay Simha Joshi (Karnataka, IN); Hozefa Yusuf Palitanawala (Foster City, CA); Pallab Rath (Bangalore, IN); Bharat Shrikrishna Paliwal (Fremont, CA); John Chaitanya Kati (Foster City, CA)
Assignee: Oracle International Corporation
G06F21/62G06F16/2456G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,584
App. No.
18/358,130
Granted
Jul 29, 2025
Kind
B2
Abstract

Techniques for automatically discovering and protecting sensitive data are disclosed. In some embodiments, a set of data objects is searched for data matching a first set of one or more regular expressions and for metadata matching a second set of one or more regular expressions. A confidence score is then generated for a particular data objects in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored in the particular data object and regular expression in the second set of one or more regular expressions that match metadata associated with the particular data object. One or more operations may be performed to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.

Claims (49)

1. A method comprising:

searching a set of data objects for data matching a first set of one or more regular expressions;

wherein the first set of one or more regular expressions include a generic regular expression and a set of one or more specific regular expressions that are associated with the generic regular expression;

generating a confidence score for a particular data object in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored by the data object;

wherein generating the confidence score comprises:

matching the generic regular expression to data in the particular data object;

responsive to matching the generic regular expression, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data;

determining whether the set of one or more specific regular expressions match data from the particular data object, and

for each specific regular expression that matches data from the particular data object, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data; and

performing at least one operation to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.

2. The method of claim 1 , wherein the confidence score is further generated as a function of contextual information associated with the data object, wherein the contextual information identifies a group category for the data object.

3. The method of claim 1 , wherein the confidence score is further generated as a function of security controls applied to the data object.

4. The method of claim 1 , wherein at least one regular expression in the first set of one or more regular expressions is associated with a first weight, wherein at least one regular expression in the first set of one or more regular expressions is associated with a second weight, and wherein the confidence score is further generated as a function of the first weight and the second weight.

5. The method of claim 1 , further comprising: identifying a sensitive data object that includes sensitive data; responsive to identifying the sensitive data object, generating at least one data regular expression and at least one metadata regular expression; wherein the at least one data regular expression is included in the first set of one or more regular expressions and the at least one metadata regular expression is included in a second set of one or more regular expressions.

6. The method of claim 1 , wherein performing the at least one operation comprises: performing at least one of: applying a data masking script to at least the particular data object to mask the sensitive data, applying a data redaction policy on a document that includes the sensitive data if the confidence score satisfies a threshold, deleting the particular data object, triggering a compliance check on the particular data object, or encrypting the sensitive data within the data object based on the confidence score.

7. The method of claim 1 , further comprising: receiving a set of discovery configurations that identify regular expressions and corresponding confidence scores; wherein the confidence score is generated by aggregating the corresponding confidence scores for regular expressions matched to data and/or metadata associated with the particular data object.

8. The method of claim 1 , further comprising: storing a plurality of sensitive data type definitions; wherein each respective sensitive data type definition defines a respective set of regular expressions for identifying data objects that satisfy the respective sensitive data type definition within a prescribed level of confidence, wherein the first set of one or more regular expressions is part of a sensitive data type definition applied to the set of data objects.

9. The method of claim 1 , further comprising: generating respective confidence scores for each data object in the set of data objects; and performing the at least one operation on a subset of data objects in the set of data objects that have confidence scores that satisfy a threshold.

10. One or more non-transitory computer-readable media storing instructions which, when executed by one or more hardware processors, cause:

searching a set of data objects for data matching a first set of one or more regular expressions;

wherein the first set of one or more regular expressions include a generic regular expression and a set of one or more specific regular expressions that are associated with the generic regular expression;

generating a confidence score for a particular data object in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored by the data object;

wherein generating the confidence score comprises:

matching the generic regular expression to data in the particular data object;

responsive to matching the generic regular expression, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data;

determining whether the set of one or more specific regular expressions match data from the particular data object, and

for each specific regular expression that matches data from the particular data object, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data; and

performing at least one operation to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.

11. The media of claim 10 , wherein the confidence score is further generated as a function of contextual information associated with the data object, wherein the contextual information identifies a group category for the data object.

12. The media of claim 10 , wherein the confidence score is further generated as a function of security controls applied to the data object.

13. The media of claim 10 , wherein at least one regular expression in the first set of one or more regular expressions is associated with a first weight, wherein at least one regular expression in the first set of one or more regular expressions is associated with a second weight, and wherein the confidence score is further generated as a function of the first weight and the second weight.

14. The media of claim 10 , wherein the instructions further cause: identifying a sensitive data object that includes sensitive data; responsive to identifying the sensitive data object, generating at least one data regular expression and at least one metadata regular expression; wherein the at least one data regular expression is included in the first set of one or more regular expressions and the at least one metadata regular expression is included in a second set of one or more regular expressions.

15. The media of claim 10 , wherein performing the at least one operation comprises: performing at least one of: applying a data masking script to at least the particular data object to mask the sensitive data, applying a data redaction policy on a document that includes the sensitive data if the confidence score satisfies a threshold, deleting the particular data object, triggering a compliance check on the particular data object, or encrypting the sensitive data within the data object based on the confidence score.

16. The media of claim 10 , wherein the instructions further cause: receiving a set of discovery configurations that identify regular expressions and corresponding confidence scores; wherein the confidence score is generated by aggregating the corresponding confidence scores for regular expressions matched to data and/or metadata associated with the particular data object.

17. The media of claim 10 , wherein the instructions further cause: storing a plurality of sensitive data type definitions; wherein each respective sensitive data type definition defines a respective set of regular expressions for identifying data objects that satisfy the respective sensitive data type definition within a prescribed level of confidence, wherein the first set of one or more regular expressions is part of a sensitive data type definition applied to the set of data objects.

18. The media of claim 10 , wherein the instructions further cause: generating respective confidence scores for each data object in the set of data objects; and performing the at least one operation on a subset of data objects in the set of data objects that have confidence scores that satisfy a threshold.

19. The media of claim 10 , further comprising: searching metadata associated with the set of data objects for metadata matching a second set of one or more regular expressions, wherein the confidence score is further generated as a function of regular expressions in the second set of one or more regular expressions that match metadata associated with the particular data object.

20. A system comprising:

one or more hardware processors;

one or more non-transitory computer-readable media storing instructions that, when executed by the one or more hardware processors, cause:

searching a set of data objects for data matching a first set of one or more regular expressions;

wherein the first set of one or more regular expressions include a generic regular expression and a set of one or more specific regular expressions that are associated with the generic regular expression;

generating a confidence score for a particular data object in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored by the data object;

wherein generating the confidence score comprises:

matching the generic regular expression to data in the particular data object;

responsive to matching the generic regular expression, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data;

determining whether the set of one or more specific regular expressions match data from the particular data object, and

for each specific regular expression that matches data from the particular data object, adjusting the confidence score to reflect a higher level of confidence that the particular data object includes sensitive data; and

performing at least one operation to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2025
From: KATI, JOHN CHAITANYA
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 071909/0123 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2023
From: PALIWAL, BHARAT SHRIKRISHNA; PALITANAWALA, HOZEFA YUSUF; RATH, PALLAB; JOSHI, VIJAY SIMHA
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064851/0773 →
Continuity (3)
Continuation 16449166 · Jun 21, 2019
Provisional Application 62748271 · Oct 19, 2018
Related Publication 20230367891A1 · Nov 16, 2023
References Cited (26)
US 8930368B2 · Furuichi · 2015 [cited by examiner]
US 11074362B2 · Conikee · 2021 [cited by examiner]
US 11500880B2 · Murray · 2022 [cited by examiner]
US 20050055369A1 · Gorelik · 2005 [cited by examiner]
US 20060095466A1 · Stevens · 2006 [cited by examiner]
US 20110173149A1 · Schon · 2011 [cited by examiner]
US 20130311456A1 · Winkler · 2013 [cited by examiner]
US 20170242876A1 · Dubost · 2017 [cited by examiner]
US 20170264619A1 · Narayanaswamy · 2017 [cited by examiner]
US 20180075104A1 · Oberbreckling · 2018 [cited by examiner]
US 20180075115A1 · Murray · 2018 [cited by examiner]
US 20180330114A1 · McGrath · 2018 [cited by examiner]
US 20190095481A1 · Lindhorst · 2019 [cited by examiner]
US 20190102438A1 · Murray · 2019 [cited by examiner]
US 20190114354A1 · Orun · 2019 [cited by examiner]
US 20190260787A1 · Zou · 2019 [cited by examiner]
US 20190370688A1 · Patel · 2019 [cited by examiner]
US 20200034473A1 · Rushan · 2020 [cited by examiner]
US 20200097587A1 · Klein · 2020 [cited by examiner]
US 20200177637A1 · Narayanaswamy · 2020 [cited by examiner]
US 20200410116A1 · Williamson · 2020 [cited by examiner]
US 20210026992A1 · Atreya · 2021 [cited by examiner]
Foreign Key Discovery, available onlie at <https://network.informatica.com/onlinehelp/analyst/961/en/index.htm#/page/data-discovery-guide/GUID-33EAF039-ECFC-49FD-96F4-A2C2A4EB857F.1.148.html>, 2 pages. [cited by applicant]
Informatica Persistent Data Masking and Data Subset User Guide (Version 9.5.0), Informatica., Part No. TDM-USG-95000-0001, Dec. 2012, 123 pages. [cited by applicant]
Jones et al., “Find Missing FK Constraints and Fixing them,” available online at <https://www.sqlservercentral.com/Forums/Topic1650014-2824-1.aspx>, Jan. 9, 2015, 12 pages. [cited by applicant]
Rostin et al., “A Machine Learning Approach to Foreign Key Discovery,” 12th International Workshop on the Web and Databases (WebDB 2009), 6 pages. [cited by applicant]
Cited By (1)
US 12,513,197