IP Library Granted Patent US 9,135,320
Granted Patent B2
US 9,135,320 · App. 13/917,263 · Granted Sep 15, 2015

System and method for data anonymization using hierarchical data clustering and perturbation

Inventors: Kanav Goyal (Panchkula, IN); Chayanika Pragya (New Delhi, IN); Rahul Garg (Ghaziabad, IN)
Assignee: Opera Solutions, LLC
G06F17/30569
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,135,320
App. No.
13/917,263
Granted
Sep 15, 2015
Kind
B2
Abstract

A system and method for data anonymization using hierarchical data clustering and perturbation is provided. The system includes a computer system and an anonymization program executed by the computer system. The system converts the data of a high-dimensional dataset to a normalized vector space and applies clustering and perturbation techniques to anonymize the data. The conversion results in each record of the dataset being converted into a normalized vector that can be compared to other vectors. The vectors are divided into disjointed, small-sized clusters using hierarchical clustering processes. Multi-level clustering can be performed using suitable algorithms at different clustering levels. The records within each cluster are then perturbed such that the statistical properties of the clusters remain unchanged.

Claims (50)

1. A system for data anonymization comprising:

a computer system for electronically receiving an original dataset and allowing a user to specify a relative importance of at least one attribute of the dataset; and

an anonymization program executed by the computer system for producing an anonymized dataset from the original dataset, the anonymization program executing:

a vector space mapping sub-process for converting each record of the original dataset to a normalized vector that can be compared to other vectors;

a hierarchical clustering sub-process for dividing the normalized vectors into disjointed k-sized groups of similar records based on a hierarchical clustering technique;

a perturbation sub-process for generating anonymized clusters from individual clusters generated by the hierarchical clustering sub-process; and

an original domain mapping sub-process to combine and remap anonymized clusters back to an original domain of the original dataset.

2. The system of claim 1 , wherein the original dataset is a high-dimensional dataset containing record-level data.

3. The system of claim 1 , wherein the vector space mapping sub-process executes a categorical to numeric conversion process.

4. The system of claim 1 , wherein the vector space mapping sub-process assigns a relative weight to every field of the records.

5. The system of claim 4 , wherein the vector space mapping sub-process normalizes the weighted records to produce the normalized vectors, which are compared with the original records to obtain mapping tables for each attribute.

6. The system of claim 1 , wherein the vector space mapping sub-process produces a normalized vector mapped dataset.

7. The system of claim 1 , wherein the hierarchical clustering sub-process divides the normalized vector mapped dataset into first-level clusters.

8. The system of claim 1 , wherein the hierarchical clustering sub-process breaks down each first-level cluster into subsequent M-th level clusters based on attributes.

9. The system of claim 1 , wherein the perturbation sub-process randomly assigns attribute values of one record to all records within each cluster.

10. The system of claim 1 , wherein the perturbation sub-process shuffles the values of an attribute among the records in each cluster by a random permutation.

11. A method for data anonymization comprising:

electronically receiving an original dataset at a computer system;

allowing a user to specify a relative importance of at least one attribute of the dataset; and

executing by the computer system an anonymization program for producing an anonymized dataset from the original dataset, the anonymization program executing:

a vector space mapping sub-process for converting each record of the original dataset to a normalized vector that can be compared to other vectors;

a hierarchical clustering sub-process for dividing the normalized vectors into disjointed k-sized groups of similar records based on a hierarchical clustering technique;

a perturbation sub-process for generating anonymized clusters from individual clusters generated by the hierarchical clustering sub-process; and

an original domain mapping sub-process to combine and remap anonymized clusters back to an original domain of the original dataset.

12. The method of claim 11 , wherein the original dataset is a high-dimensional dataset containing record-level data.

13. The method of claim 11 , wherein the vector space mapping sub-process executes a categorical to numeric conversion process.

14. The method of claim 11 , wherein the vector space mapping sub-process assigns a relative weight to every field of the records.

15. The method of claim 14 , wherein the vector space mapping sub-process normalizes the weighted records to produce the normalized vectors, which are compared with the original records to obtain mapping tables for each attribute.

16. The method of claim 11 , wherein the vector space mapping sub-process produces a normalized vector mapped dataset.

17. The method of claim 11 , wherein the hierarchical clustering sub-process divides the normalized vector mapped dataset into first-level clusters.

18. The method of claim 11 , wherein the hierarchical clustering sub-process breaks down each first-level cluster into subsequent M-th level clusters based on attributes.

19. The method of claim 11 , wherein the perturbation sub-process randomly assigns attribute values of one record to all records within each cluster.

20. The method of claim 11 , wherein the perturbation sub-process shuffles the values of an attribute among the records in each cluster by a random permutation.

21. A computer-readable medium having computer-readable instructions stored thereon which, when executed by a computer system, cause the computer system to perform the steps of:

electronically receiving an original dataset at a computer system;

allowing a user to specify a relative importance of at least one attribute of the dataset; and

executing by the computer system an anonymization program for producing an anonymized dataset from the original dataset, the anonymization program executing:

a vector space mapping sub-process for converting each record of the original dataset to a normalized vector that can be compared to other vectors;

a hierarchical clustering sub-process for dividing the normalized vectors into disjointed k-sized groups of similar records based on a hierarchical clustering technique;

a perturbation sub-process for generating anonymized clusters from individual clusters generated by the hierarchical clustering sub-process; and

an original domain mapping sub-process to combine and remap anonymized clusters back to an original domain of the original dataset.

22. The computer-readable medium of claim 21 , wherein the original dataset is a high-dimensional dataset containing record-level data.

23. The computer-readable medium of claim 21 , wherein the vector space mapping sub-process executes a categorical to numeric conversion process.

24. The computer-readable medium of claim 21 , wherein the vector space mapping sub-process assigns a relative weight to every field of the records.

25. The computer-readable medium of claim 24 , wherein the vector space mapping sub-process normalizes the weighted records to produce the normalized vectors, which are compared with the original records to obtain mapping tables for each attribute.

26. The computer-readable medium of claim 21 , wherein the vector space mapping sub-process produces a normalized vector mapped dataset.

27. The computer-readable medium of claim 21 , wherein the hierarchical clustering sub-process divides the normalized vector mapped dataset into first-level clusters.

28. The computer-readable medium of claim 21 , wherein the hierarchical clustering sub-process breaks down each first-level cluster into subsequent M-th level clusters based on attributes.

29. The computer-readable medium of claim 21 , wherein the perturbation sub-process randomly assigns attribute values of one record to all records within each cluster.

30. The computer-readable medium of claim 21 , wherein the perturbation sub-process shuffles the values of an attribute among the records in each cluster by a random permutation.

Assignments (9)
CHANGE OF NAME Recorded Aug 2, 2021
From: OPERA SOLUTIONS OPCO, LLC
To: ELECTRIFAI, LLC
Reel/Frame 057047/0300 →
TRANSFER STATEMENT AND ASSIGNMENT Recorded Oct 21, 2018
From: WHITE OAK GLOBAL ADVISORS, LLC
To: OPERA SOLUTIONS OPCO, LLC
Reel/Frame 047276/0107 →
SECURITY AGREEMENT Recorded Jul 7, 2016
From: OPERA SOLUTIONS USA, LLC; OPERA SOLUTIONS, LLC; OPERA SOLUTIONS GOVERNMENT SERVICES, LLC; BIQ, LLC; LEXINGTON ANALYTICS INCORPORATED; OPERA PAN ASIA LLC
To: WHITE OAK GLOBAL ADVISORS, LLC
Reel/Frame 039277/0318 →
TERMINATION AND RELEASE OF IP SECURITY AGREEMENT Recorded Jul 7, 2016
From: PACIFIC WESTERN BANK, AS SUCCESSOR IN INTEREST BY MERGER TO SQUARE 1 BANK
To: OPERA SOLUTIONS, LLC
Reel/Frame 039277/0480 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2016
From: OPERA SOLUTIONS, LLC
To: OPERA SOLUTIONS U.S.A., LLC
Reel/Frame 039089/0761 →
SECURITY INTEREST Recorded Dec 7, 2015
From: OPERA SOLUTIONS, LLC
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 037243/0788 →
SECURITY INTEREST Recorded Feb 9, 2015
From: OPERA SOLUTIONS, LLC
To: SQUARE 1 BANK
Reel/Frame 034923/0238 →
SECURITY INTEREST Recorded Nov 21, 2014
From: OPERA SOLUTIONS, LLC
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 034311/0552 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2013
From: GOYAL, KANAV; PRAGYA, CHAYANIKA; GARG, RAHUL
To: OPERA SOLUTIONS, LLC
Reel/Frame 031330/0514 →
Continuity (2)
Provisional Application 61659178 · Jun 13, 2012
Related Publication 20130339359A1 · Dec 19, 2013