IP Library › Granted Patent US 11,694,118
Granted Patent B2
US 11,694,118 · App. 17/093,563 · Granted Jul 4, 2023

System and method for data visualization using machine learning and automatic insight of outliers associated with a set of data

Inventors: Ashish Mittal (Foster City, CA); Victor Belyaev (San Jose, CA); Steve Simon Joseph Fernandez (Columbia, MO); Gabby Rubin (Sunnyvale, CA); Alextair Mascarenhas (Foster City, CA); Samar Lotia (Cupertino, CA); Alvin Raj (Woburn, MA); John Fuller (Chicago, IL); Saugata Chowdhury (Sunnyvale, CA)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06N20/00G06F3/0481G06F3/0486G06F16/2272G06F16/248G06F16/252G06F16/26G06T11/206G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,694,118
App. No.
17/093,563
Granted
Jul 4, 2023
Kind
B2
Abstract

In accordance with various embodiments, described herein are systems and methods for use of computer-implemented machine learning to automatically determine insights of facts, segments, outliers, or other information associated with a set of data, for use in generating visualizations of the data. In accordance with an embodiment, the system can use a machine learning process to automatically determine one or more outliers or findings within the data, based on, for example, determining a plurality of combinations representing pairs of attribute dimensions within a data set, from which a general explanation or pattern can be determined for one or more attributes, and then comparing particular values for attributes, with the determined pattern for those attributes. Information describing such outliers or findings can be graphically displayed at a user interface, as text, graphs, charts, or other types of visualizations, and used as a starting point for further analysis of the data set.

Claims (43)

1. A system for use of machine learning in a data visualization environment, to automatically determine, for a set of data, one or more outliers or findings within the data set, comprising:

one or more computer systems or devices, including a microprocessor, and a data visualization service executing thereon that provides access to a database having a data set associated with a plurality of attributes and dimensions;

wherein the data visualization service provides access by a client system to communicate requests for information associated with the data set, and receive at a user interface, findings associated with the data set and displayed as visualizations within the user interface;

wherein the data visualization service is adapted to:

receive, from the client system, an indication of a target attribute of interest as provided within the data set;

determine a plurality of dimension attributes associated with the data set;

for each of one or more pairs of the dimension attributes, calculate, for a first dimension attribute, expected values of the target attribute with respect to values of the target attribute associated with a second dimension attribute;

determine observed values for the target attribute within the data set;

generate ranked findings associated with the data set based on a comparison of the expected values and observed values for the target attribute; and

report the ranked findings to the client system, for use in generating a data visualization for initial display at the user interface;

wherein, after reporting the ranked findings to the client system, the data visualization service automatically continues to evaluate pairs of dimension attributes to determine additional findings, and report the additional findings as the ranked findings to the client system for display at the user interface, the additional findings displayed replacing at least one of the ranked findings in the initial display.

2. The system of claim 1 , wherein the findings are ranked according to relative deviation between the expected values and the observed values.

3. The system of claim 1 , wherein the data visualization service is provided within a cloud computing environment, and receives requests from client systems to display visualizations of the findings associated with the data set.

4. The system of claim 1 , wherein the data visualization environment operates to:

for each pair of dimension attributes, operate a regression process to calculate an expected value of a target attribute for those dimension attributes with respect to each distinctive value in the other dimension attribute;

generate a finding for each dimension attribute in the pair of attribute dimensions, based on discrepancies between expected and observed values.

5. The system of claim 1 , wherein the data visualization service is adapted to automatically determine a data visualization algorithm for use in generating one or more visualizations.

6. A method for automatically determining and reporting findings associated with data within a database or data set, comprising:

providing a data visualization service that communicates with and provides access to a database having a data set associated with a plurality of attributes and dimensions;

providing access by a client system to communicate requests for information associated with the data set, and receive at a user interface, findings associated with the data set and displayed as visualizations within the user interface;

receiving, from the client system, an indication of a target attribute of interest as provided within the data set;

determining a plurality of dimension attributes associated with the data set;

for each of one or more pairs of the dimension attributes, calculating, for a first dimension attribute, expected values of the target attribute with respect to values of the target attribute associated with a second dimension attribute;

determining observed values for the target attribute within the data set;

generating ranked findings associated with the data set based on a comparison of the expected values and observed values for the target attribute; and

reporting the ranked findings to the client system, for use in generating a data visualization for initial display at the user interface;

wherein the data visualization service automatically continues to evaluate pairs of dimension attributes to determine additional findings, and report the additional findings as the ranked findings to the client system for display at the user interface, the additional findings displayed replacing at least one of the ranked findings in the initial display.

7. The method of claim 6 , wherein the findings are ranked according to relative deviation between the expected values and the observed values.

8. The method of claim 6 , wherein the data visualization service is provided within a cloud computing environment, and receives requests from client systems to display visualizations of the findings associated with the data set.

9. The method of claim 6 , wherein the data visualization environment operates to:

for each pair of dimension attributes, operate a regression process to calculate an expected value of a target attribute for those dimension attributes with respect to each distinctive value in the other dimension attribute;

generate a finding for each dimension attribute in the pair of attribute dimensions, based on discrepancies between expected and observed values.

10. The method of claim 6 , wherein the data visualization service is adapted to automatically determine a data visualization algorithm for use in generating one or more visualizations.

11. A non-transitory computer readable medium, including instructions stored thereon which when read and executed by one or more computers including one or more processors cause the one or more computers to perform a method comprising:

providing a data visualization service that communicates with and provides access to a database having a data set associated with a plurality of attributes and dimensions;

providing access by a client system to communicate requests for information associated with the data set, and receive at a user interface, findings associated with the data set and displayed as visualizations within the user interface;

receiving, from the client system, an indication of a target attribute of interest as provided within the data set;

determining a plurality of dimension attributes associated with the data set;

for each of one or more pairs of the dimension attributes, calculating, for a first dimension attribute, expected values of the target attribute with respect to values of the target attribute associated with a second dimension attribute;

determining observed values for the target attribute within the data set;

generating ranked findings associated with the data set based on a comparison of the expected values and observed values for the target attribute; and

reporting the ranked findings to the client system, for use in generating a data visualization for initial display at the user interface;

wherein the data visualization service automatically continues to evaluate pairs of dimension attributes to determine additional findings, and report the additional findings as the ranked findings to the client system for display at the user interface, the additional findings displayed replacing at least one of the ranked findings in the initial display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2021
From: MITTAL, ASHISH; BELYAEV, VICTOR; FERNANDEZ, STEVE SIMON JOSEPH; RUBIN, GABBY; MASCARENHAS, ALEXTAIR; LOTIA, SAMAR; RAJ, ALVIN; FULLER, JOHN; CHOWDHURY, SAUGATA
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 054832/0858 →
Continuity (6)
Continuation 16148680 · Oct 1, 2018
Provisional Application 62566271 · Sep 29, 2017
Provisional Application 62566264 · Sep 29, 2017
Provisional Application 62566263 · Sep 29, 2017
Provisional Application 62566265 · Sep 29, 2017
Related Publication 20210073682A1 · Mar 11, 2021