IP Library › Granted Patent US 10,049,128
Granted Patent B1
US 10,049,128 · App. 14/588,054 · Granted Aug 14, 2018

Outlier detection in databases

Inventor: Yuting Zhang (Sichuan, CN)
Assignee: Symantec Corporation
G06F17/30371G06F17/30289G06F19/704
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,049,128
App. No.
14/588,054
Granted
Aug 14, 2018
Kind
B1
Abstract

Various systems, methods, and processes for identifying outliers in a data set stored in a database are disclosed. A subset of data is extracted from a data set. Data descriptors are allocated to the subset of data. A model of the subset of data is created based on attributes of the data descriptors. An iteration of an outlier detection process based on the model is then executed. The outlier detection process evaluates the subset of data, and the outlier detection process evaluates the data set based on the results of the evaluation of the subset of data. The outlier detection process, which can implement and/or use a Random Sample Consensus (RANSAC) algorithm, identifies outliers in the data set stored in the database.

Claims (71)

1. A method comprising:

extracting a first subset of data from a data set stored in a database, wherein the data set is a string data set having at least one outlier string data member that reduces the quality of the database storing the string data set;

creating a first model for the first subset of data using one or more string descriptors describing attributes of string data, wherein the one or more string descriptors are selected from: string length, string character set, co-occurrence of string elements, frequency of string character appearance, entropy of the string data set, similarity of string data, and segmentation of subset of data, the first model being created by:

allocating one or more string descriptors to the first subset of data;

calculating a value distribution of features for one or more allocated string descriptors;

generating a fingerprint for each allocated string descriptor of the first subset of data based on the value distribution, the fingerprint presenting a data distribution feature of an allocated string descriptor; and

generating the first model based on the fingerprint;

executing a first iteration of an outlier detection process based on the first model, the first model being evaluated responsive to the first iteration of the outlier detection process to determine whether the first model identifies an optimum number of outliers in the first subset of data, wherein

the outlier detection process evaluates the data set based on results of an evaluation of the first subset of data, and

the outlier detection process is a multi-iteration Random Sample Consensus (RANSAC) algorithm based on the one or more string descriptors; and

responsive to determining that an outlier detection threshold has been met based on the first iteration, identifying the at least one outlier string data member in the string data set based on the executing step; and

removing the at least one outlier string data member from the string data set, thereby improving the quality of the database storing the string data set.

2. The method of claim 1 , further comprising:

executing a clustering process to identify a first filtering threshold of the first subset of data that is distinct from a similarity threshold and is based on the first model; and

filtering the data set based on the first filtering threshold of the first subset of data.

3. The method of claim 2 , further comprising:

executing a second iteration of the outlier detection process, wherein executing the second iteration creates a second model and identifies the similarity threshold for a second subset of data that is distinct from the first filtering threshold and is based on the second model, wherein the second subset of data is part of the data set.

4. The method of claim 3 , wherein identifying the at least one outlier string data member in the data set further comprises:

comparing the first model of the first subset of data and the second model of the second subset of data, and

filtering the data set using a model with the higher similarity threshold, the higher similarity threshold being based on a comparison of the distribution of the one or more string descriptors.

5. The method of claim 2 , further comprising:

using the clustering process, creating the fingerprint, the fingerprint associating the first subset of data and the data set based on at least one of the string descriptors.

6. The method of claim 5 , further comprising:

determining a value distribution of the string descriptors based on the fingerprint created using the clustering process, and

calculating a similarity between the first subset of data and the data set based on the value distribution.

7. The method of claim 1 , wherein the data set is part of a database application, and identifying the at least one outlier string data member in the data set further comprises detecting dirty data in the data set that is part of the database application.

8. A computer readable storage medium comprising program instructions executable to:

extract a first subset of data from a data set stored in a database, wherein the data set is a string data set having at least one outlier string data member that reduces the quality of the database storing the string data set;

create a first model for the first subset of data using one or more string descriptors describing attributes of string data, wherein the one or more string descriptors are selected from: string length, string character set, co-occurrence of string elements, frequency of string character appearance, entropy of the string data set, similarity of string data, and segmentation of subset of data, the first model being created by:

allocating one or more string descriptors to the first subset of data;

calculating a value distribution of features for one or more allocated string descriptors;

generating a fingerprint for each allocated string descriptor for the subset of data based on the value distribution, the fingerprint presenting a data distribution feature of an allocated string descriptor; and

generating the first model based on the fingerprint;

execute a first iteration of an outlier detection process based on the first model, the first model being evaluated responsive to the first iteration of the outlier detection process to determine whether the first model identifies an optimum number of outliers in the first subset of data, wherein

the outlier detection process evaluates the string data set based on results of an evaluation of the first subset of data, and

the outlier detection process is a multi-iteration Random Sample Consensus (RANSAC) algorithm based on the one or more string descriptors; and

responsive to determining that an outlier detection threshold has been met based on the first iteration, identify the at least one outlier string data member in the string data set based on the executing step; and

remove the at least one outlier string data member from the string data set, thereby improving the quality of the database storing the string data set.

9. The computer readable storage medium of claim 8 , further comprising:

executing a clustering process to identify a first filtering threshold of the first subset of data that is distinct from a similarity threshold and is based on the first model,

filtering the data set based on the first filtering threshold of the first subset of data, and

executing a second iteration of the outlier detection process, wherein executing the second iteration creates a second model and identifies the similarity threshold for a second subset of data that is distinct from the first filtering threshold and is based on the second model, wherein the second subset of data is part of the data set.

10. The computer readable storage medium of claim 9 , wherein identifying the at least one outlier string data member in the data set further comprises:

comparing the first model of the first subset of data and the second model of the second subset of data, and

filtering the data set using a model with the higher similarity threshold, the higher similarity threshold being based on a comparison of the distribution of the one or more string descriptors.

11. The computer readable storage medium of claim 9 , further comprising:

using the clustering process, creating the fingerprint, the fingerprint associating the first subset of data and the data set based on at least one of the string descriptors,

determining a value distribution of the string descriptors based on the fingerprint created using the clustering process, and

calculating a similarity between the first subset of data and the data set based on the value distribution.

12. The computer readable storage medium of claim 8 , wherein the data set is part of a database application, and identifying the at least one outlier string data member in the data set further comprises detecting dirty data in the data set that is part of the database application.

13. A system comprising:

one or more processors; and

a memory coupled to the one or more processors, wherein the memory stores program instructions executable by the one or more processors to:

extract a first subset of data from a data set stored in a database, wherein the data set is a string data set having at least one outlier string data member that reduces the quality of the database storing the string data set;

create a first model for the first subset of data using one or more string descriptors describing attributes of string data, wherein the one or more string descriptors are selected from: string length, string character set, co-occurrence of string elements, frequency of string character appearance, entropy of the string data set, similarity of string data, and segmentation of subset of data, the first model being created by:

allocating one or more string descriptors to the first subset of data;

calculating a value distribution of features for one or more allocated string descriptors;

generating a fingerprint for each allocated string descriptor for the subset of data based on the value distribution, the fingerprint presenting a data distribution feature of an allocated string descriptor; and

generating the first model based on the fingerprint;

execute a first iteration of an outlier detection process based on the first model, the first model being evaluated responsive to the first iteration of the outlier detection process to determine whether the first model identifies an optimum number of outliers in the first subset of data, wherein

the outlier detection process evaluates the string data set based on results of an evaluation of the first subset of data, and

the outlier detection process is a multi-iteration Random Sample Consensus (RANSAC) algorithm based on the one or more string descriptors; and

responsive to determining that an outlier detection threshold has been met based on the first iteration, identify the at least one outlier string data member in the string data set based on the executing step; and

remove the at least one outlier string data member from the string data set, thereby improving the quality of the database storing the string data set.

14. The system of claim 13 , further comprising:

executing a clustering process to identify a first filtering threshold of the first subset of data that is distinct from a similarity threshold and is based on the first model,

filtering the data set based on the first filtering threshold of the first subset of data, and

executing a second iteration of the outlier detection process, wherein executing the second iteration creates a second model and identifies the similarity threshold for a second subset of data that is distinct from the first filtering threshold and is based on the second model, wherein the second subset of data is part of the data set.

15. The system of claim 14 , wherein identifying the at least one outlier string data member in the data set further comprises:

comparing the first model of the first subset of data and the second model of the second subset of data, and

filtering the data set using a model with the higher similarity threshold, the higher similarity threshold being based on a comparison of the distribution of the one or more string descriptors.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2019
From: SYMANTEC CORPORATION
To: CA, INC.
Reel/Frame 051144/0918 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2015
From: ZHANG, YUTING
To: SYMANTEC CORPORATION
Reel/Frame 034611/0338 →
Cited By (1)
US 12,748,119