IP Library › Granted Patent US 11,113,256
Granted Patent B2
US 11,113,256 · App. 16/455,133 · Granted Sep 7, 2021

Automated data discovery with external knowledge bases

Inventors: Lingtao Zhang (Coquitlam, CA); Chang Lu (Vancouver, CA); Amit Kumar (Fremont, CA)
Assignee: salesforce.com, inc.
G06F16/221G06F16/285G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,113,256
App. No.
16/455,133
Granted
Sep 7, 2021
Kind
B2
Abstract

System and methods are described for improving automated data discovery analysis in a cloud computing environment. A method includes receiving a request to analyze a data set stored in the memory device, the data set including one or more columns, the one or more columns including one or more data values in one or more cells of each column; classify each of the one or more columns as a type of column; for a selected one of the one or more columns, if the selected column's type is an external type, join one or more columns of an external knowledge base correlated to the selected column into the data set to create an expanded data set; and execute an automated data discovery model on the expanded data set.

Claims (44)

1. A computing system, comprising:

a processing device; and

a memory device coupled to the processing device, the memory device having instructions stored thereon that, in response to execution by the processing device, cause the processing device to:

receive a request to analyze a data set stored in the memory device, the data set including one or more columns, the one or more columns including one or more data values in one or more cells of each column;

classify each of the one or more columns as a type of column;

for a selected one of the one or more columns, if the selected column's type is an external type, join one or more columns of an external knowledge base correlated to the selected column into the data set to create an expanded data set;

execute an automated data discovery model on the expanded data set to automatically determine insights from the expanded data set not discoverable by executing the automated data discovery model on the data set; and

train the automated data discovery model on the expanded data set.

2. The computing system of claim 1 , wherein the memory device having instructions stored thereon that, in response to execution by the processing device, cause the processing device to:

classify the one or more columns by, for each column, determining a type of data stored in the column and storing the type for the column;

wherein a type is one of internal, external and unknown.

3. The computing system of claim 2 , wherein the memory device having instructions stored thereon that, in response to execution by the processing device, cause the processing device to:

for a selected one of the one or more columns, if the selected column's type is internal, join one or more replicated columns of the data set correlated to the selected column into the data set to create the expanded data set.

4. The computing system of claim 2 , wherein the data set includes at least one column having an internal type comprising data that is managed by a user, and the external knowledge base includes at least one column having an external type comprising data publicly available over a computer network.

5. A computer-implemented method comprising:

receiving a request to analyze a data set stored in the memory device, the data set including one or more columns, the one or more columns including one or more data values in one or more cells of each column;

classifying each of the one or more columns as a type of column;

for a selected one of the one or more columns, if the selected column's type is an external type, joining one or more columns of an external knowledge base correlated to the selected column into the data set to create an expanded data set;

executing an automated data discovery model on the expanded data set to automatically determine insights from the expanded data set not discoverable by executing the automated data discovery model on the data set and

train the automated data discovery model on the expanded data set.

6. The computer-implemented method of claim 5 , comprising classifying the one or more columns by, for each column, determining a type of data stored in the column and storing the type for the column; wherein a type is one of internal, external and unknown.

7. The computer-implemented method of claim 6 , comprising for a selected one of the one or more columns, if the selected column's type is internal, joining one or more replicated columns of the data set correlated to the selected column into the data set to create the expanded data set.

8. The computer-implemented method of claim 6 , wherein the data set includes at least one column having an internal type comprising data that is managed by a user, and the external knowledge base includes at least one column having an external type comprising data publicly available over a computer network.

9. A tangible, non-transitory computer-readable storage medium having instructions encoded thereon which, when executed by a processing device, cause the processing device to:

receive a request to analyze a data set stored in the memory device, the data set including one or more columns, the one or more columns including one or more data values in one or more cells of each column;

classify each of the one or more columns as a type of column;

for a selected one of the one or more columns, if the selected column's type is an external type, join one or more columns of an external knowledge base correlated to the selected column into the data set to create an expanded data set;

execute an automated data discovery model on the expanded data set to automatically determine insights from the expanded data set not discoverable by executing the automated data discovery model on the data set; and

train the automated data discovery model on the expanded data set.

10. The tangible, non-transitory computer-readable storage medium of claim 9 , having instructions encoded thereon which, when executed by a processing device, cause the processing device to:

classify the one or more columns by, for each column, determining a type of data stored in the column and storing the type for the column;

wherein a type is one of internal, external and unknown.

11. The tangible, non-transitory computer-readable storage medium of claim 10 , having instructions encoded thereon which, when executed by a processing device, cause the processing device to:

for a selected one of the one or more columns, if the selected column's type is internal, join one or more replicated columns of the data set correlated to the selected column into the data set to create the expanded data set.

12. The tangible, non-transitory computer-readable storage medium of claim 10 , wherein the data set includes at least one column having an internal type comprising data that is managed by a user, and the external knowledge base includes at least one column having an external type comprising data publicly available over a computer network.

13. A processing system comprising:

a database including a data set, the data set including one or more columns, the one or more columns including one or more data values in one or more cells of each column; and

an analyzer to receive a request to analyze the data set, the analyzer including:

a classifier to classify each of the one or more columns as a type of column; and

a join engine to, for a selected one of the one or more columns, if the selected column's type is an external type, join one or more columns of an external knowledge base correlated to the selected column into the data set to create an expanded data set;

an analysis engine to execute an automated data discovery model on the expanded data set to automatically determine insights from the expanded data set not discoverable by executing the automated data discovery model on the data set; and

a model trainer to train the automated data discovery model on the expanded data set.

14. The processing system of claim 13 , wherein the classifier to classify the one or more columns by, for each column, determining a type of data stored in the column and storing the type for the column; wherein a type is one of internal, external and unknown.

15. The processing system of claim 14 , wherein the join engine to for a selected one of the one or more columns, if the selected column's type is internal, join one or more replicated columns of the data set correlated to the selected column into the data set to create the expanded data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2019
From: ZHANG, LINGTAO; LU, CHANG; KUMAR, AMIT
To: SALESFORCE.COM, INC.
Reel/Frame 049750/0241 →
Continuity (1)
Related Publication 20200409919A1 · Dec 31, 2020