IP Library Granted Patent US 8,306,932
Granted Patent B2
US 8,306,932 · App. 12/384,776 · Granted Nov 6, 2012

System and method for adaptive data masking

Assignee: Infosys Technologies Limited
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,306,932
App. No.
12/384,776
Granted
Nov 6, 2012
Kind
B2
Abstract

A method for adaptive data masking of a database is provided. The method comprises extracting data from a first database and providing one or more predefined rules for masking the extracted data. The method further comprises masking a first portion of extracted data using a trained Artificial Neural Network (ANN), where the ANN is trained for masking at least one database having properties similar to the first database. The masked and unmasked data is aggregated to arrive at an output structurally similar to the extracted data. The method furthermore comprises determining a deviation value between the arrived output and expected output of the extracted data, and adapting the trained ANN automatically according to data masking requirements of the first database, if the deviation value is more than a predefined value.

Claims (42)

1. A method for adaptive data masking of data present in a first database, the method comprising:

extracting data from the first database;

applying one or more predefined rules based on properties of the first database for masking the extracted data, wherein the one or more predefined rules are designed based on association between data organized in the first database, and further wherein the one or more predefined rules specify a portion of the extracted data required to be masked;

segregating a first portion of the extracted data required for masking from a second portion of the extracted data not required for masking based on the one or more predefined rules;

selectively masking the first portion of the extracted data using a trained Artificial Neural Network (ANN) based on the one or more predefined rules, wherein the ANN is previously trained for masking at least one database having properties similar to the first database;

aggregating the masked data and the second portion of the extracted data to arrive at an output structurally similar to the extracted data;

checking if the output is in accordance with the one or more predefined rules;

modifying the output based on the one or more predefined rules, if the output is not in accordance with the one or more predefined rules; and

adapting the trained ANN automatically according to the modified output.

2. The method of claim 1 , wherein the trained ANN selectively masks the first portion of the extracted data based on a weight matrix, the weight matrix being predetermined based on the one or more predefined rules.

3. The method of claim 1 further comprising storing the arrived output in an intermediate database.

4. The method of claim 3 further comprising labeling the intermediate database as a final mask database if the output is in accordance with the one or more predefined rules.

5. The method of claim 1 , wherein adapting the trained ANN automatically according to the modified output comprises using one or more training algorithms for neural networks.

6. The method of claim 1 , wherein the extracted data may be a datasheet comprising ‘m’ rows and ‘n’ columns or any combination thereof.

7. The method of claim 1 , wherein the one or more predefined rules are represented as metadata, the metadata being a sparse way of representing the extracted data.

8. A system for adaptive data masking, the system comprising:

a first database;

a data extractor configured to extract data from the first database;

a data segregator configured to segregate a first portion of the extracted data required for masking from a second portion of the extracted data, not required for masking using one or more predefined rules designed based on properties of the first database, wherein the one or more predefined rules specify a portion of the extracted data required to be masked, and further wherein the one or more predefined rules are based on association between data organized in the first database;

an adaptive Artificial Neural Network (ANN) configured to selectively mask the first portion of the extracted data based on the one or more predefined rules, wherein the ANN is previously trained for masking at least one database having properties similar to the first database;

an aggregator configured to aggregate the masked data and the second portion of the extracted data, to arrive at an output data;

an intermediate database configured to store the arrived output data; and

a quality checker configured to:

check if the output data is in accordance with the one or more predefined rules; and

modify the output data based on the one or more predefined rules, if the output is not in accordance with the one or more predefined rules, wherein the trained ANN adapts itself automatically according to the modified output.

9. The system of claim 8 , wherein the intermediate database is labeled as a final mask database if the output data is in accordance with the one or more predefined rules.

10. The system of claim 8 , wherein the quality checker is configured to store the one or more predefined rules.

11. A computer program product comprising a non-transitory computer usable medium having a computer readable program code embodied therein for adaptive data masking of data present in a first database, the computer program product comprising:

program instructions for extracting data from the first database;

program instructions for applying one or more predefined rules based on properties of the first database for masking the extracted data, wherein the one or more predefined rules are designed based on association between data organized in the first database, and further wherein the one or more predefined rules specify a portion of the extracted data required to be masked;

program instructions for segregating a first portion of the extracted data required for masking from a second portion of the extracted data, not required for masking based on the one or more predefined rules;

program instructions for selectively masking the first portion of the extracted data using a trained Artificial Neural Network (ANN), wherein the ANN is previously trained for masking at least one database having properties similar to the first database;

program instructions for aggregating the masked data and the second portion of the extracted data to arrive at an output structurally similar to the extracted data;

program instructions for checking if the output is in accordance with the one or more predefined rules;

program instructions for modifying the output based on the one or more predefined rules, if the output is not in accordance with the one or more predefined rules; and

program instructions for adapting the trained ANN automatically according to the modified output.

12. The computer program product of claim 11 , wherein the trained ANN selectively masks the first portion of the extracted data based on a weight matrix, the weight matrix being predetermined based on the one or more predefined rules.

13. The computer program product of claim 11 further comprising program instructions for storing the arrived output in an intermediate database.

14. The computer program product of claim 13 further comprising program instructions for labeling the intermediate database as a final mask database if the output is in accordance with the one or more predefined rules.

15. The computer program product of claim 11 , wherein the program instructions for adapting the trained ANN automatically according to the modified output uses one or more training algorithms for neural networks.

16. The computer program product of claim 11 , wherein the extracted data may be a datasheet comprising ‘m’ rows and ‘n’ columns or any combination thereof.

17. The computer program product of claim 11 , wherein the one or more predefined rules are represented as metadata, the metadata being a sparse way of representing the extracted data.

Assignments (2)
CHANGE OF NAME Recorded Mar 18, 2013
From: INFOSYS TECHNOLOGIES LIMITED
To: INFOSYS LIMITED
Reel/Frame 030050/0683 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2009
From: SAXENA, ASHUTOSH; GUJJARY, VISHAL ANJAIAH; SURNI, KUMAR
To: INFOSYS TECHNOLOGIES LIMITED
Reel/Frame 023022/0070 →
Priority Claims (1)
IN 882/CHE/2008 · Apr 8, 2008 · national
Continuity (1)
Related Publication 20090281974A1 · Nov 12, 2009