IP Library › Granted Patent US 11,163,745
Granted Patent B2
US 11,163,745 · App. 16/650,610 · Granted Nov 2, 2021

Statistical fingerprinting of large structure datasets

Inventors: Arthur Coleman (Carmel Valley, CA); Tsz Ling Christina Leung (Foster City, CA); Martin Rose (Superior, CO); Chivon Powers (Burlingame, CA); Natarajan Shankar (Pleasanton, CA)
Assignee: LiveRamp, Inc.
G06F16/2282G06F16/221G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,163,745
App. No.
16/650,610
Granted
Nov 2, 2021
Kind
B2
Abstract

A system and method for statistical fingerprinting of structured datasets begins by dividing the structured database into groups of data subsets. These subsets are created based on the structure of the data; for example, data delineated by columns and rows may be broken into subsets by designating each column as a subset. A fingerprint is derived from each subset, and then the fingerprint for each subset is combined in order to create an overall fingerprint for the dataset. By applying this process to a “wild file” of unknown provenance, and comparing the result to a data owner's files, it may be determined if data in the wild file was wrongfully acquired from the data owner.

Claims (32)

1. A method of fingerprinting for dynamic structured databases, the method comprising the steps of:

dividing a structured database into a plurality of subsets, wherein each of the plurality of subsets comprise columns in the structured database;

deriving a fingerprint for each of the plurality of subsets;

combining the fingerprint for each of the plurality of subsets to create a first fingerprint for the structured database;

applying a first timestamp to the first fingerprint for the structured database;

at a later time after changes occur in the structured database, again dividing the structured database into a plurality of subsets and deriving the fingerprint for each of the plurality of subsets, then combining the fingerprints for each of the plurality of subsets to create a second fingerprint for the structured database;

applying a second timestamp to the second fingerprint for the structured database; and

performing a time dimension analysis using the first fingerprint and second fingerprint in order to account for time drift within the structured database.

2. The method of claim 1 , further comprising the step of profiling each of the plurality of columnar data sets by data type.

3. The method of claim 2 , further comprising the step of pre-processing each of the plurality of columnar data sets.

4. The method of claim 3 , further comprising the step of applying at least one of a plurality of statistical tests to the plurality of columnar data sets.

5. The method of claim 4 , wherein at least one of the columnar data sets comprises a quantitative data set.

6. The method of claim 5 , wherein the at least one of a plurality of statistical tests applied to the quantitative data set is selected from the set consisting of Mean, Median, Mode, Min, Max, Standard Deviation and Variance.

7. The method of claim 4 , wherein at least one of the columnar data sets comprises a qualitative data set.

8. The method of claim 7 , wherein at least one of the plurality of statistical tests applied to the plurality of columnar data sets is selected from the set consisting of two-sample Chi Square, Chi Square Goodness of Fit, and Chi Square Test of Independence.

9. The method of claim 1 , further comprising the step of comparing the first fingerprint for the structured database or the second fingerprint for the structured database or both the first fingerprint and second fingerprint for the structured database to a data owner fingerprint for a data owner structured database to determine if the structured database was derived from the data owner structured database.

10. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause it to:

divide a structured database into a plurality of subsets, wherein each of the plurality of subsets comprise columns in the structured database;

derive a fingerprint for each of the plurality of subsets;

combine the fingerprint for each of the plurality of subsets to create a fingerprint for the structured database;

apply a first timestamp to the first fingerprint for the structured database;

at a later time after changes occur in the structured database, again divide the structured database into a plurality of subsets and derive the fingerprint for each of the plurality of subsets, then combine the fingerprints for each of the plurality of subsets to create a second fingerprint for the structured database;

apply a second timestamp to the second fingerprint for the structured database; and

perform a time dimension analysis using the first fingerprint and second fingerprint in order to account for time drift within the structured database.

11. The non-transitory computer-readable storage medium of claim 10 , further comprising stored instructions that, when executed by a computer, cause it to profile each of the plurality of columnar data sets by data type.

12. The non-transitory computer-readable storage medium of claim 11 , further comprising stored instructions that, when executed by a computer, cause it to pre-process each of the plurality of columnar data sets.

13. The non-transitory computer-readable storage medium of claim 12 , further comprising stored instructions that, when executed by a computer, cause it to apply at least one of a plurality of statistical tests to the plurality of columnar data sets.

14. The non-transitory computer-readable storage medium of claim 13 , wherein at least one of the columnar data sets comprises a quantitative data set.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the at least one of a plurality of statistical tests applied to the quantitative data set is selected from the set consisting of Mean, Median, Mode, Min, Max, Standard Deviation and Variance.

16. The non-transitory computer-readable storage medium of claim 13 , wherein at least one of the columnar data sets comprises a qualitative data set.

17. The non-transitory computer-readable storage medium of claim 16 , wherein at least one of the plurality of statistical tests applied to the plurality of columnar data sets is selected from the set consisting of two-sample Chi Square, Chi Square Goodness of Fit, and Chi Square Test of Independence.

18. The non-transitory computer-readable storage medium of claim 10 , further comprising stored instructions that, when executed by a computer, cause it to compare the first fingerprint for the structured database or the second fingerprint for the structured database or both the first fingerprint and second fingerprint for the structured database to a fingerprint for a data owner structured database to determine if the structured database was derived from the data owner structured database.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2021
From: COLEMAN, ARTHUR; LEUNG, TSZ LING CHRISTINA; ROSE, MARTIN; POWERS, CHIVON; SHANKAR, NATARAJAN
To: ACXIOM LLC
Reel/Frame 055146/0045 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2021
From: ACXIOM LLC
To: LIVERAMP, INC.
Reel/Frame 055148/0774 →
Continuity (2)
Provisional Application 62568720 · Oct 5, 2017
Related Publication 20210200735A1 · Jul 1, 2021
Cited By (3)
US 12,314,263 US 12,675,484 US 12,748,757