IP Library Granted Patent US 11,423,045
Granted Patent B2
US 11,423,045 · App. 16/401,202 · Granted Aug 23, 2022

Augmented analytics techniques for generating data visualizations and actionable insights

Inventors: Gopinath Rajendiran (Tiruvannamalai, IN); Aaradhana Sridharan (Chennai, IN); Vidhyadharan Deivamani (Theni, IN); Vijay Anand Chidambaram (Chennai, IN); Ulrich Kalex (Potsdam, DE)
Assignee: SOFTWARE AG
G06F16/26G06F3/0484G06F16/248G06T11/206
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,045
App. No.
16/401,202
Granted
Aug 23, 2022
Kind
B2
Abstract

A data analysis system is provided. Processing resources are configured to at least: identify features within a dataset, identify potential features of interest therefrom, and enable selection of one of the identified potential features of interest. Responsive to an identified potential feature of interest being selected: (a) algorithms are run on the dataset to identify at least one related feature that the selected feature of interest is most likely and/or most heavily influenced by; (b) a display is generated to include a visual representation of each related feature, each including associated data value representations; and (c) a visual representation can be selected. A data value representation is selectable together with the selected visual representation. Responsive selection of the visual representation, (a)-(c) are repeated. Responsive to a data value representation being selected in (c), the dataset is filtered based on it, and the repetition is performed with the filtered dataset.

Claims (77)

1. A data analysis system, comprising:

an electronic interface over which a dataset is accessible; and

processing resources including at least one processor and a memory coupled thereto, the processing resources being configured to execute instructions stored to the memory to at least:

access the dataset using the electronic interface;

identify features within the dataset, wherein different features describe different properties of or pertaining to one or more data elements in the dataset;

identify potential features of interest from the identified features;

enable selection of one of the identified potential features of interest; and

responsive to a selection of one of the identified potential features of interest:

(a) select, based on a data type that is associated with the selected feature of interest, at least one of a plurality of algorithms and run the selected at least one of the plurality of algorithms on the dataset to identify at least one related feature that satisfies an influence threshold with respect to the selected feature of interest, wherein the at least one of the plurality of algorithms that are selected are variable across successive repetitions;

(b) generate a display including a visual representation of each related feature, each visual representation including representations of data values associated with the respective related feature, wherein the representations of the data values form parts of the respective visual representations;

(c) enable selection of one of the representations of the data values from one of the displayed visual representations;

(d) responsive to a representation of a data value being selected in (c), filter rows from the dataset based on the selected representation of the data value and the respective related feature for the one of the representations of the data values that is selected; and

repeat at least (a) (c), wherein the repetition is performed in connection with the filtered dataset, wherein the dataset filtration is maintained through successive repetitions, and wherein features not previously identified as being related in an earlier instance of (a) are identifiable in successive repetitions.

2. The system of claim 1 , wherein the identification of the potential features of interest from the identified features includes, for each of the identified features, calculating the respective identified feature's variance.

3. The system of claim 2 , wherein the identification of the potential features of interest from the identified features further includes normalizing the variances and designating a predetermined number of the identified features having the highest normalized variances as the potential features of interest.

4. The system of claim 2 , wherein the identification of the potential features of interest from the identified features further includes designating a predetermined number of the identified features having the highest variances as the potential features of interest.

5. The system of claim 4 , wherein the predetermined number is user-configurable.

6. The system of claim 1 , wherein the generation of the display including the visual representation of each related feature includes:

determining a chart type for each related feature; and

forming each visual representation in accordance with the determined chart type for the respective related feature.

7. The system of claim 6 , wherein the chart type for each related feature is determined automatically and is changeable in response to user input.

8. The system of claim 1 , wherein the generation of the display including the visual representation for each related feature includes creating a bar chart for each related feature, and wherein the representations of data values are bars in the bar charts.

9. The system of claim 1 , wherein the algorithms are run so as to identify, as a related feature, each related feature for which a corresponding visual representation was selected in a previous repetition.

10. The system of claim 1 , wherein the display is generated to include a visual representation for each related feature for which a corresponding visual representation was selected in a previous repetition.

11. The system of claim 1 , wherein the algorithms are run on a common set of the identified features across each repetition, regardless of whether any identified features have been identified as related features, and wherein different algorithms are runnable across successive repetitions.

12. The system of claim 1 , wherein:

the processing resources are further configured to execute instructions stored to the memory to at least determine which one of a plurality of classes of algorithms is to be run; and

the algorithms that are selected to run for each iteration include algorithms included in the determined class.

13. The system of claim 12 , wherein the determination of which one of the plurality of classes of algorithms is to be run is based on the data type of the selected feature of interest, wherein the class of algorithms to be run is variable across successive repetitions.

14. The system of claim 1 , wherein the identification of the at least one related feature that satisfies the influence threshold includes:

determining a predictive value of each algorithm run; and

identifying, from the algorithm determined to have the highest predictive value, those features(s) that satisfy the influence threshold with respect to the selected feature of interest, as the at least one related feature.

15. The system of claim 1 , wherein a plurality of related features is identified.

16. The system of claim 1 , wherein the processing resources are further configured to respond to the number of features returned in (a) falling below a threshold by performing (b) and preventing (c) and (d).

17. The system of claim 1 , wherein the processing resources are further configured to respond to the number of features returned in (a) falling below a threshold by performing (b) and preventing (c) and (d) in response to a majority of the features returned in (a) having a relevance-related score less than a predetermined value.

18. The system of claim 1 , wherein (a) further comprises:

determining a score for each feature; and

selecting as the at least one related feature that satisfies the influence threshold; (i) all related features having an associated score above a predetermined threshold, (ii) the related features having the highest scores, up to N related features being selected, or (iii) the related features having the highest scores, up to M % related features being selected.

19. A method for analyzing data in a dataset, the method comprising:

accessing the dataset using an electronic interface;

identifying features within the dataset;

identifying potential features of interest from the identified features;

enabling selection of one of the identified potential features of interest; and

responsive to a selection of one of the identified potential features of interest:

(a) select, based on a data type that is associated with the selected feature of interest, at least one of a plurality of computer-implemented algorithms and run the selected at least one of the plurality of algorithms on the dataset to identify at least one related feature that satisfies an influence threshold with respect to the selected feature of interest, wherein the at least one of the plurality of algorithms that are selected are variable across successive repetitions;

(b) causing a display to be generated, the display including a visual representation of each related feature, each visual representation including representations of data values associated with the respective related feature, wherein the representations of the data values form parts of the respective visual representations;

(c) enabling selection of one of the representations of the data values from one of the displayed visual representations;

(d) responsive to a representation of a data value being selected in (c), filter rows from the dataset based on the selected representation of the data value and the respective related feature for the one of the representations of the data values that is selected; and

repeating at least (a)-(c), wherein the repetition is performed in connection with the filtered dataset, wherein the dataset filtration is maintained through successive repetitions, and wherein features not previously identified as being related in an earlier instance of (a) are identifiable in successive repetitions.

20. The method of claim 19 , wherein the identification of the potential features of interest from the identified features includes:

for each of the identified features, calculating the respective identified feature's variance; and

designating a predetermined number of the identified features having the highest normalized or non-normalized variances as the potential features of interest.

21. The method of claim 19 , further comprising:

determining a chart type for each related feature; and

forming each visual representation in accordance with the determined chart type for the respective related feature;

wherein the chart type for each related feature is determined automatically and is changeable in response to user input.

22. The method of claim 19 , wherein the algorithms are run on a common set of the identified features across each repetition, regardless of whether any identified features have been identified as related features.

23. The method of claim 19 , further comprising determining which one of a plurality of classes of algorithms is to be run;

wherein the algorithms that are selected to run for each iteration include algorithms included in the determined class.

24. The method of claim 19 , wherein the identification of the at least one related feature that satisfies the influence threshold includes:

determining a predictive value of each algorithm run; and

identifying, from the algorithm determined to have the highest predictive value, the feature(s) that satisfy the influence threshold with respect to the selected feature of interest, as the at least one related feature.

25. The method of claim 19 , further comprising, in response to the number of features returned in (a) falling below a threshold, (i) generating a user prompt, and (ii) performing (b) and preventing (c) and (d).

26. The method of claim 25 , wherein (i) and (ii) are performed in response to a majority of the features returned in (a) having a relevance-related score less than a predetermined value.

27. A non-transitory computer readable storage medium storing instructions that, when executed by a computer including at least one hardware processor, control the computer to at least:

access an electronic dataset;

identify features within the dataset, wherein different features describe different properties of or pertaining to one or more data elements in the dataset;

identify potential features of interest from the identified features;

enable selection of one of the identified potential features of interest; and

responsive to a selection of one of the identified potential features of interest:

(a) select, based on a data type that is associated with the selected feature of interest, at least one of a plurality of computer-implemented algorithms and run the selected at least one of the plurality of algorithms on the dataset to identify a plurality of related features that satisfies an influence threshold with respect to the selected feature of interest, wherein the at least one of the plurality of algorithms that are selected are variable across successive repetitions;

(b) cause a display to be generated, the display including a visual representation of each related feature, each visual representation including representations of data values associated with the respective related feature, wherein the representations of the data values form parts of the respective visual representations;

(c) enable selection of one of the representations of the data values from one of the displayed visual representations;

(d) responsive to a representation of a data value being selected in (c), filter rows from the dataset based on the selected representation of the data value and the respective related feature for the one of the representations of the data values that is selected; and

repeat at least (a)-(c), wherein the repetition is performed in connection with the filtered dataset,

wherein the algorithms are run on a common set of the identified features across each repetition, regardless of whether any identified features have been identified as related features, and

in response to the number of features returned in (a) falling below a threshold, perform (b) and preventing (c) and (d).

Assignments (4)
CHANGE OF NAME Recorded Dec 9, 2024
From: MOSEL BIDCO AG
To: SOFTWARE GMBH
Reel/Frame 069548/0271 →
MERGER Recorded Dec 9, 2024
From: SOFTWARE AG
To: MOSEL BIDCO AG
Reel/Frame 069548/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2024
From: SOFTWARE GMBH
To: SAG ALFABET GMBH
Reel/Frame 070069/0859 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2019
From: RAJENDIRAN, GOPINATH; SRIDHARAN, AARADHANA; DEIVAMANI, VIDHYADHARAN; CHIDAMBARAM, VIJAY ANAND; KALEX, ULRICH
To: SOFTWARE AG
Reel/Frame 049066/0415 →
Continuity (1)
Related Publication 20200349170A1 · Nov 5, 2020