IP Library › Granted Patent US 11,188,845
Granted Patent B2
US 11,188,845 · App. 16/148,675 · Granted Nov 30, 2021

System and method for data visualization using machine learning and automatic insight of segments associated with a set of data

Inventors: Victor Belyaev (San Jose, CA); Gabby Rubin (Sunnyvale, CA); Samar Lotia (Cupertino, CA); Alvin Raj (Woburn, MA); John Fuller (Chicago, IL)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06N20/00G06F3/0481G06F3/0486G06F16/2272G06F16/248G06F16/252G06F16/26G06T11/206G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,188,845
App. No.
16/148,675
Granted
Nov 30, 2021
Kind
B2
Abstract

In accordance with various embodiments, described herein are systems and methods for use of computer-implemented machine learning to automatically determine insights of facts, segments, outliers, or other information associated with a set of data, for use in generating visualizations of the data. In accordance with an embodiment, the system can use a machine learning process to automatically determine one or more segments within a data set, associated with a target attribute value, based on, for example, the use of a classification and regression tree and a combination of different driving factors, or same driving factors with different values. Information describing segments associated with the data set can be graphically displayed at a user interface, as text, graphs, charts, or other types of visualizations, and used as a starting point for further analysis of the data set.

Claims (55)

1. A system for use of machine learning in a data visualization environment, to automatically determine, for a set of data, one or more segments associated with a target attribute, the system comprising:

one or more computer systems or devices, including a microprocessor, and a data visualization cloud service executing thereon that provides access to a data set; and

wherein the data visualization cloud service receives a request from a client application, to provide an explanation of segments associated with a target attribute within the data set, wherein the target attribute is identified in said request;

wherein the data visualization cloud service, responsive to receiving said request, operates according to a machine learning algorithm to:

retrieve a plurality of attributes from the data set, each attribute comprising a column of data values;

remove high-cardinality attributes within the retrieved plurality of attributes;

remove non-correlated attributes from the retrieved plurality of attributes that have no correlation with the target attribute, leaving only correlated attributes from the retrieved plurality of attributes that are correlated with the target attribute;

remove duplicate correlated attributes having high correlation with other correlated attributes from the retrieved plurality of attributes;

construct a decision tree by recursively splitting data values remaining in the data set after removing attributes that have no correlation with the target attribute and removing correlated attribute having high correlation with other correlated attributes; and

generate, from the decision tree, a plurality of segments associated with the target attribute for display; and

provide, to the client application, as a response to said request, data describing the plurality of segments from the data set and associated with the target attribute identified in said request, that can be graphically displayed as visualizations of the data set.

2. The system of claim 1 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

3. The system of claim 1 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

4. The system of claim 1 , wherein the data visualization cloud service operates to:

construct the decision tree by recursively splitting the prepared data set according to one or more other attributes; and

generate a plurality of segments having leaf nodes associated with observation information from the decision tree.

5. The system of claim 1 , wherein:

the data visualization cloud service, responsive to receiving said request further operates according to a machine learning algorithm to associate each of the plurality of segments with a plurality of driving factors for the target attribute.

6. A method for use of machine learning in a data visualization environment, to automatically determine, for a set of data, one or more segments associated with a target attribute, the method comprising:

providing at one or more computer systems or devices, including a microprocessor, a data visualization cloud service executing thereon that provides access to a data set; and

receiving, at the data visualization cloud service, a request from a client application, to provide an explanation of segments associated with a target attribute within the data set, wherein the target attribute is identified in said request;

responsive to receiving said request from the client application:

retrieving a plurality of attributes from the data set, each attribute comprising a column of data values,

removing high-cardinality attributes within the retrieved plurality of attributes,

removing non-correlated attributes from the retrieved plurality of attributes that have no correlation with the target attribute leaving only correlated attributes from the retrieved plurality of attributes that are correlated with the target attribute,

removing duplicate correlated attributes having high correlation with other correlated attributes from the retrieved plurality of attributes,

constructing a decision tree by recursively splitting data values remaining in the data set after removing attributes that have no correlation with the target attribute and removing correlated attribute having high correlation with other correlated attributes, and

generating, from the decision tree, a plurality of segments associated with the target attribute for display, and

providing, to the client application, as a response to said request, data describing the plurality of segments associated with the data set and associated with the target attribute identified in said request, that can be graphically displayed as visualizations of the data set.

7. The method of claim 6 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

8. The method of claim 6 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

9. The method of claim 6 , wherein the data visualization cloud service operates to:

construct the decision tree by recursively splitting the prepared data set according to one or more other attributes; and

generate a plurality of segments having leaf nodes associated with observation information from the decision tree.

10. The method of claim 6 , further comprising:

associating each of the plurality of segments with a description and a plurality of driving factors for the target attribute.

11. A non-transitory computer-readable storage medium including instructions stored thereon for supporting use of machine learning in a data visualization environment to automatically determine for a set of data one or more segments associated with a target attribute, which instructions, when read and executed by one or more computers cause the one or more computers to perform a method comprising:

receiving, at a data visualization cloud service that provides access to a data set, a request from a client application, to provide an explanation of segments associated with a target attribute within the data set, wherein the target attribute is identified in said request;

responsive to receiving said request:

retrieving a plurality of attributes from the data set, each attribute comprising a column of data values,

removing high-cardinality attributes within the retrieved plurality of attributes, leaving only correlated attributes from the retrieved plurality of attributes that are correlated with the target attribute,

removing non-correlated attributes from the retrieved plurality of attributes that have no correlation with the target attribute,

removing duplicate correlated attributes having high correlation with other correlated attributes from the retrieved plurality of attributes,

constructing a decision tree by recursively splitting data values remaining in the data set after removing attributes that have no correlation with the target attribute and removing correlated attribute having high correlation with other correlated attributes, and

generating, from the decision tree, a plurality of segments associated with the target attribute for display, and

providing, to the client application, as a response to said request, data describing the plurality of segments from the data set and associated with the target attribute identified in said request, that can be graphically displayed as visualizations of the data set.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

13. The non-transitory computer-readable storage medium of claim 11 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

14. The non-transitory computer-readable storage medium of claim 11 , wherein the data visualization cloud service operates to:

construct the decision tree by recursively splitting the prepared data set according to one or more other attributes; and

generate a plurality of segments having leaf nodes associated with observation information from the decision tree.

15. The non-transitory computer-readable storage medium of claim 11 , wherein the instructions further cause the data visualization cloud service to:

associate each of the plurality of segments with a plurality of driving factors for the target attribute.

16. The non-transitory computer-readable storage medium of claim 11 , wherein the instructions further cause the data visualization cloud service to:

associate each of the plurality of segments with a description and a plurality of driving factors for the target attribute.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2018
From: RUBIN, GABBY
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 047073/0176 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2018
From: BELYAEV, VICTOR; LOTIA, SAMAR; RAJ, ALVIN; FULLER, JOHN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 047073/0179 →
Continuity (5)
Provisional Application 62566263 · Sep 29, 2017
Provisional Application 62566264 · Sep 29, 2017
Provisional Application 62566265 · Sep 29, 2017
Provisional Application 62566271 · Sep 29, 2017
Related Publication 20190102703A1 · Apr 4, 2019