IP Library Granted Patent US 11,423,009
Granted Patent B2
US 11,423,009 · App. 16/880,411 · Granted Aug 23, 2022

System and method to prevent formation of dark data

Inventors: Brendan Stennett (Toronto, CA); Bryan Smith (Toronto, CA); Yousuf Chowdhary (Toronto, CA)
Assignee: ThinkData Works, Inc.
G06F16/2365G06F16/2228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,009
App. No.
16/880,411
Granted
Aug 23, 2022
Kind
B2
Abstract

A method is provided for preventing dark data in a data set. At a time t 1 , a first version of the data set is received. The first version is analyzed and its parameters are gathered in a first statistical profile. The first statistical profile is stored. At a time t 2 , a second version of the data set is received. The second version is analyzed and its parameters are gathered in a second statistical profile. The second statistical profile is stored. The first and second statistical profiles are compared and a similarity index is created. If the similarity index exceeds a pre-set threshold, dissimilarity is flagged and a responsive action is taken.

Claims (30)

1. A server for preventing formation of dark data in a data set, the server having a processor, a communication subsystem, and a memory, each in communication with the processor, the memory storing instructions which when executed by the processor configure the server to:

at a time t 1 , receive a first version of the data set to be ingested into a data warehouse;

analyze the first version and gather its parameters in a first statistical profile;

store the first statistical profile as a suggested schema, while the first version of the data set is automatically ingested into the data warehouse;

at a time t 2 , receive a second version of the data set to be ingested into the data warehouse;

analyze the second version and gather its parameters in a second statistical profile based on the suggested schema;

store the second statistical profile;

compare each parameter of the first and second statistical profiles with an upper or lower limit and create a similarity index;

wherein the comparison is assessed against a tolerance level and dissimilarity is flagged if the parameter's distance from the upper or lower limit is outside the tolerance level; and

wherein the tolerance level is based on observed patterns from statistical profiles of historical data sets previously ingested; and

if the similarity index exceeds a pre-set threshold, flag dissimilarity and take a responsive action; and

automatically halt the ingestion of the second version of the data set until the dissimilarity flag is addressed through a responsive action.

2. The server of claim 1 , wherein the responsive action includes performing further automated analysis of at least one of the first or second versions.

3. The server of claim 1 , wherein the responsive action includes sending a notification to an owner or administrator of the first or second versions.

4. The server of claim 3 , wherein the responsive action further includes allowing the owner or administrator to correct the version of the data set or substitute the version of the data set with a corrected version.

5. The server of claim 4 , wherein the correction is performed automatically upon a prompt.

6. The server of claim 1 , wherein the first version of the data set is a reference version.

7. The server of claim 1 , wherein if the similarity index is below the pre-set threshold, ingestion of the second version is allowed to proceed.

8. The server of claim 7 , wherein the second version is merged with at least one other previously ingested version of the data set.

9. The server of claim 7 , wherein the second version is saved as a reference version for future comparison.

10. The server of claim 1 , wherein the statistical profile includes, for each parameter of the data set, at least one of: average, minimum, maximum, standard deviation, variance, number of unique, null percentage, coefficient of variation, frequency.

11. The server of claim 1 , wherein the statistical profile is stored in association with at least one of: dataset ID, revision, update timestamp.

12. The server of claim 1 , wherein the statistical profile is stored in a relational database.

13. The server of claim 1 , wherein the statistical profile is stored as a JSON object.

14. The server of claim 1 , wherein the parameter is a text string and dissimilarity is flagged if the string length is above or below a character limit.

15. The server of claim 1 , wherein the method is repeated as newer versions of the data set are received.

16. The server of claim 1 , wherein the method is repeated as previously received versions of the data set are ingested.

17. The server of claim 1 , wherein the method is repeated at scheduled intervals regardless of whether the previously received versions were ingested.

18. The server of claim 1 , wherein the data sets are from multiple geographic locations.

19. The server of claim 1 , wherein the first and second versions are from multiple geographic locations.

Assignments (2)
SECURITY INTEREST Recorded Mar 19, 2021
From: THINKDATA WORKS INC.
To: THE BANK OF NOVA SCOTIA
Reel/Frame 055652/0915 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2020
From: STENNETT, BRENDAN; SMITH, BRYAN; CHOWDHARY, YOUSUF
To: THINKDATA WORKS INC.
Reel/Frame 053498/0920 →
Continuity (2)
Provisional Application 62853907 · May 29, 2019
Related Publication 20200379974A1 · Dec 3, 2020