IP Library Granted Patent US 11,599,667
Granted Patent B1
US 11,599,667 · App. 16/990,809 · Granted Mar 7, 2023

Efficient statistical techniques for detecting sensitive data

Inventors: Aurelian Tutuianu (Iasi, RO); Daniel Voinea (Iasi, RO); Petru-Serban Cehan (Iasi, RO); Silviu Catalin Poede (Iasi, RO); Adrian Cadar (Iasi, RO); Marian-Razvan Udrea (Iasi, RO); Brent Gregory (Iasi, RO)
Assignee: Amazon Technologies, Inc.
G06F21/6245G06F16/2462G06F16/29G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,599,667
App. No.
16/990,809
Filed
Aug 11, 2020
Granted
Mar 7, 2023
Kind
B1
Examiner
CHEN, CAI Y
Art Unit
2425
USPC
726/26
Abstract

A candidate attribute combination of a first data set is identified, such that the candidate attribute combination meets a data type similarity criterion with respect to a collection of data types of sensitive information for which the first data set is to be analyzed. A collection of input features is generated for a machine learning model from the candidate attribute combination, including at least one feature indicative of a statistical relationship between the values of the candidate attribute combination and a second data set. An indication of a predicted probability of a presence of sensitive information in the first data set is obtained using the machine learning model.

Claims (55)

1. A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices cause the one or more computing devices to:

obtain a first data set indicating human population density as a function of geographical location;

identify a plurality of structured data objects to be analyzed for a presence of geographical location details pertaining to individuals, wherein the geographical location details are expressed using a plurality of numeric data types, wherein individual ones of the structured data objects comprise a plurality of records, and wherein individual ones of plurality of records comprise values of a plurality of attributes;

select a sample of records from a particular structured data object of the plurality of structured data objects;

identify one or more candidate attribute combinations from the plurality of attributes of the records of the sample, wherein (a) individual ones of the candidate attribute combinations meet a data type similarity criterion with respect to the plurality of numeric data types and (b) attribute values of individual ones of the candidate attribute combinations satisfy one or semantic filtration criteria associated with geographical location details;

generate, corresponding to individual ones of the one or more candidate attribute combinations and the first data set, a collection of input features for a classification model, including at least one feature indicative of a statistical relationship between human population density and attribute values of the candidate attribute combinations; and

transmit an indication of a probability of a presence of geographical location details in the particular structured data object, wherein the probability is obtained from the classification model using at least the collection of input features.

2. The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

obfuscate, by applying one or more transformation operations, raw values of one or more attributes of the plurality of attributes, such that obfuscated versions of the raw values are used to generate the collection of input features.

3. The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

obtain, via one or more programmatic interfaces, a request for sensitive data presence analysis of one or more data stores, including a data store at which the plurality of structured data objects is stored, wherein the plurality of structured data objects is identified in response to the request for sensitive data presence analysis.

4. The system as recited in claim 1 , wherein to identify one or more candidate attribute combinations from the plurality of attributes, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

compute a ratio of (a) a number of records of the sample whose attribute values for a particular attribute lie within a valid range of values corresponding to a particular representation of a geographic location and (b) a number of records of the sample whose attribute values for the particular attribute are non-empty.

5. The system as recited in claim 1 , wherein to identify one or more candidate attribute combinations from the plurality of attributes, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

compute a similarity metric associated with respective values of a first attribute of the plurality of attributes and a second attribute of the plurality of attributes.

6. A computer-implemented method, comprising:

obtaining a first data set indicating a distribution of one or more properties of a group of entities with respect to which targeted information presence analysis is to be performed;

identifying one or more candidate attribute combinations from a plurality of attributes of records of a second data set, wherein individual ones of the candidate attribute combinations meet a data type similarity criterion with respect to a collection of data types of targeted information of entities of the group of entities;

generating, corresponding to individual ones of the one or more candidate attribute combinations and the first data set, a collection of input features for a machine learning model, including at least one feature indicative of a statistical relationship between the distribution of the one or more properties and attribute values of an individual candidate attribute combination; and

obtaining an indication of a predicted probability of a presence of targeted information in the second data set, wherein the predicted probability is obtained from the machine learning model using at least the collection of input features.

7. The computer-implemented method as recited in claim 6 , wherein the one or more properties of the group of entities comprise a population density as a function of geographical location.

8. The computer-implemented method as recited in claim 6 , further comprising:

obtaining an indication, via one or more programmatic interfaces, that a data store is to be analyzed for presence of targeted information; and

selecting a subset of the data store in response to obtaining the indication, wherein the subset comprises the second data set.

9. The computer-implemented method as recited in claim 6 , further comprising:

determining at least a portion of a topology of a data store comprising the second data set; and

selecting the second data set from the data store based at least in part on the topology.

10. The computer-implemented method as recited in claim 6 , wherein identifying the one or more candidate attribute combinations from the plurality of attributes of records comprises:

applying one or more semantic filters to values of individual attributes of the plurality of attributes, wherein the one or more semantic filters are defined based at least in part on characteristics of the targeted information.

11. The computer-implemented method as recited in claim 6 , wherein the first data set comprises a plurality of data points, and wherein generating the collection of input features comprises:

assigning individual data points of the plurality of data points to respective cells of a grid; and

assigning respective values of a particular candidate attribute combination to respective cells of the grid.

12. The computer-implemented method as recited in claim 11 , wherein generating the collection of input features comprises:

identifying a group of cells of the grid for which counts of assigned values of the particular candidate attribute combination exceed zero; and

determining a metric of correlation between (a) the counts of assigned values of the particular candidate attribute combination of individual cells of the group of cells and (b) values obtained from the data points of the first data set which were assigned to individual cells of the group of cells.

13. The computer-implemented method as recited in claim 6 , wherein generating the collection of input features comprises:

applying one or more smoothing functions to values of a particular candidate attribute combination, wherein the one or more smoothing functions comprise one or more of (a) a Laplacian smoothing function or (b) a Gaussian kernel density smoothing function.

14. The computer-implemented method as recited in claim 6 , wherein generating the collection of input features comprises:

determining a metric of correlation between values of at least a pair of attributes of a particular candidate attribute combination.

15. The computer-implemented method as recited in claim 6 , wherein generating the collection of input features comprises:

executing a goodness-of-fit test, wherein a collection of expected values of the goodness-of-fit test is based at least in part on values of the first data set, and wherein a collection of observed values of the goodness-of-fit test is based at least in part on the values of a particular candidate attribute combination.

16. One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:

identify one or more candidate attribute combinations from a plurality of attributes of records of a first data set, wherein individual ones of the candidate attribute combinations meet a data type similarity criterion with respect to a collection of data types of targeted information of entities of a group of entities;

generate, corresponding to individual ones of the one or more candidate attribute combinations, a collection of input features for a machine learning model, including at least one feature indicative of a statistical relationship between (a) attribute values of the individual candidate attribute combination and (b) a second data set indicating one or more properties of a group of entities with respect to which targeted information presence analysis is to be performed; and

obtain an indication of a predicted probability of a presence of targeted information in the first data set, wherein the predicted probability is obtained from the machine learning model using at least the collection of input features.

17. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the machine learning model comprises one or more of: (a) a logistic regression model or (b) a neural network-based model.

18. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

in response to determining that the predicted probability exceeds a threshold, transmit, via one or more programmatic interfaces, one or more of: (a) an indication that the predicted probability exceeds the threshold, (b) names of attributes of a particular candidate attribute combination, wherein the predicted probability was generated as output by the machine learning model for one or more input features corresponding to the particular candidate attribute combination or (c) an indication of a cluster of records identified from the first data set, wherein individual records of the cluster satisfy a structural similarity criterion, and wherein the predicted probability was generated as output by the machine learning model for one or more input features corresponding to an attribute combination of the cluster.

19. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

obtain, via one or more programmatic interfaces, an indication of one or more of: (a) a hyper-parameter of the machine learning model, (b) a feature generation algorithm, (c) a technique to be used to identify candidate attribute combinations, or (d) a record sampling algorithm.

20. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the targeted information comprises one or more of: (a) global positioning system (GPS) coordinates, (b) account identifiers, or (c) user identifiers.

21. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

subsequent to obtaining the indication of the predicted probability, initiating one or more of: (a) a notification, (b) a submission of a request to an owner of the first data set to verify presence of the targeted information in the first data set, (c) a submission of a request to an owner of the first data set to approve an administrative action with respect to the first data set, (d) isolation of at least a portion of the first data set, (e) encryption of at least a portion of the first data set, or (f) a deletion of at least a portion of the first data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2020
From: TUTUIANU, AURELIAN; VOINEA, DANIEL; CEHAN, PETRU-SERBAN; POEDE, SILVIU CATALIN; CADAR, ADRIAN; UDREA, MARIAN-RAZVAN; GREGORY, BRENT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053473/0897 →
Cited By (12)
US 12,216,682 US 12,259,847 US 12,271,911 US 12,287,782 US 12,314,425 US 12,326,949 US 12,353,432 US 12,452,690 US 12,547,615 US 12,579,294 US 12,602,907 US 12,682,108