IP Library › Granted Patent US 10,832,171
Granted Patent B2
US 10,832,171 · App. 16/148,680 · Granted Nov 10, 2020

System and method for data visualization using machine learning and automatic insight of outliers associated with a set of data

Inventors: Ashish Mittal (Foster City, CA); Victor Belyaev (San Jose, CA); Steve Simon Joseph Fernandez (Columbia, MO); Gabby Rubin (Sunnyvale, CA); Alextair Mascarenhas (Foster City, CA); Samar Lotia (Cupertino, CA); Alvin Raj (Woburn, MA); John Fuller (Chicago, IL); Saugata Chowdhury (Sunnyvale, CA)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06N20/00G06F3/0481G06F3/0486G06F16/2272G06F16/248G06F16/252G06F16/26G06T11/206G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,832,171
App. No.
16/148,680
Filed
Oct 1, 2018
Granted
Nov 10, 2020
Kind
B2
Art Unit
2611
USPC
345/440
Abstract

In accordance with various embodiments, described herein are systems and methods for use of computer-implemented machine learning to automatically determine insights of facts, segments, outliers, or other information associated with a set of data, for use in generating visualizations of the data. In accordance with an embodiment, the system can use a machine learning process to automatically determine one or more outliers or findings within the data, based on, for example, determining a plurality of combinations representing pairs of attribute dimensions within a data set, from which a general explanation or pattern can be determined for one or more attributes, and then comparing particular values for attributes, with the determined pattern for those attributes. Information describing such outliers or findings can be graphically displayed at a user interface, as text, graphs, charts, or other types of visualizations, and used as a starting point for further analysis of the data set.

Claims (41)

1. A system for use of machine learning in a data visualization environment, to automatically determine, for a set of data, one or more outliers or findings within the data set, comprising:

one or more computer systems or devices, including a microprocessor, and a data visualization cloud service executing thereon that provides access to a data set associated with a plurality of attribute dimensions; and

wherein the data visualization environment, responsive to receiving a request from a client application, to provide an explanation of outliers or findings within the data set, operates according to a machine learning algorithm to:

determine a plurality of combinations representing pairs of attribute dimensions associated with the data set, from which a general explanation or pattern can be determined for one or more attributes;

compare particular values for the one or more attributes, with the determined pattern for those attributes, including a process of:

for each pair of dimension attributes, calculating, for a first dimension attribute, expected values of a target attribute with respect to values associated with a second dimension attribute;

generating outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between the expected values and observed values; and

provide, to the client application, data describing outliers or findings for the one or more attributes associated the data set which outliers or findings are graphically displayed as visualizations of the data set, and continuing the process for additional pairs of dimension attributes, and providing an indication of additional outliers for display at the user interface.

2. The system of claim 1 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

3. The system of claim 1 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

4. The system of claim 1 , wherein the data visualization environment operates to:

for each pair of dimension attributes, use linear regression to calculate an expected value of the target attribute for those dimension attributes with respect to each distinctive value in the other dimension attribute;

generate a list of outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between expected and observed values; and

surface a set of outlier information.

5. A method for use of machine learning in a data visualization environment, to automatically determine, for a set of data, one or more outliers or findings within the data set, comprising:

providing at one or more computer systems or devices, including a microprocessor, a data visualization cloud service executing thereon that provides access to a data set associated with a plurality of attribute dimensions; and

responsive to receiving a request from a client application, to provide an explanation of outliers or findings within the data set:

determining a plurality of combinations representing pairs of attribute dimensions associated with the data set, from which a general explanation or pattern can be determined for one or more attributes;

comparing particular values for the one or more attributes, with the determined pattern for those attributes, including a process of:

for each pair of dimension attributes, calculating, for a first dimension attribute, expected values of a target attribute with respect to values associated with a second dimension attribute;

generating outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between the expected values and observed values; and

providing, to the client application, data describing outliers or findings for the one or more attributes associated the data set which outliers or findings are graphically displayed as visualizations of the data set, and continuing the process for additional pairs of dimension attributes, and providing an indication of additional outliers for display at the user interface.

6. The method of claim 5 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

7. The method of claim 5 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

8. The method of claim 5 , wherein the data visualization environment operates to:

for each pair of dimension attributes, use linear regression to calculate an expected value of the target attribute for those dimension attributes with respect to each distinctive value in the other dimension attribute;

generate a list of outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between expected and observed values; and

surface a set of outlier information.

9. A non-transitory computer-readable storage medium including instructions stored thereon which when read and executed by one or more computers cause the one or more computers to perform a method comprising:

responsive to receiving, at a data visualization cloud service that provides access to a data set associated with a plurality of attribute dimensions, a request from a client application, to provide an explanation of outliers or findings within the data set:

determining a plurality of combinations representing pairs of attribute dimensions associated with the data set, from which a general explanation or pattern can be determined for one or more attributes;

comparing particular values for the one or more attributes, with the determined pattern for those attributes, including a process of:

for each pair of dimension attributes, calculating, for a first dimension attribute, expected values of a target attribute with respect to values associated with a second dimension attribute;

generating outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between the expected values and observed values; and

providing, to the client application, data describing outliers or findings for the one or more attributes associated the data set which outliers or findings are graphically displayed as visualizations of the data set, and continuing the process for additional pairs of dimension attributes, and providing an indication of additional outliers for display at the user interface.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the data visualization cloud service is provided within a cloud computing environment, and receives requests from the client application at a client computer system or device, to display visualizations of the data set at the client application.

11. The non-transitory computer-readable storage medium of claim 9 , wherein a user interface at the client application includes a data attribute panel that enables the client application to display a data set, and to enable drag and drop of attributes to a canvas in the user interface, for use in creating visualizations.

12. The non-transitory computer-readable storage medium of claim 9 , wherein the data visualization environment operates to:

for each pair of dimension attributes, use linear regression to calculate an expected value of the target attribute for those dimension attributes with respect to each distinctive value in the other dimension attribute;

generate a list of outliers for each dimension attribute in the pair of attribute dimensions, based on discrepancies between expected and observed values; and

surface a set of outlier information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2018
From: FERNANDEZ, STEVE SIMON JOSEPH; RUBIN, GABBY
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 047073/0199 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2018
From: MITTAL, ASHISH; BELYAEV, VICTOR; MASCARENHAS, ALEXTAIR; LOTIA, SAMAR; RAJ, ALVIN; FULLER, JOHN; CHOWDHURY, SAUGATA
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 047073/0206 →
Continuity (5)
Provisional Application 62566263 · Sep 29, 2017
Provisional Application 62566264 · Sep 29, 2017
Provisional Application 62566265 · Sep 29, 2017
Provisional Application 62566271 · Sep 29, 2017
Related Publication 20190102921A1 · Apr 4, 2019
Cited By (1)
US 12,443,844