IP Library Granted Patent US 11,675,765
Granted Patent B2
US 11,675,765 · App. 17/329,519 · Granted Jun 13, 2023

Top contributor recommendation for cloud analytics

Inventors: Ying Wu (Dublin, IE); Malte Christian Kaufmann (Dublin, IE); Alan McShane (Raheny, GB); Anirban Banerjee (Kilcullen, IE); Gareth Maguire (Newbridge, IE)
Assignee: BUSINESS OBJECTS SOFTWARE LTD.
G06F16/2237G06F16/2264G06F18/2113G06F18/2321G06F18/23213
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,765
App. No.
17/329,519
Granted
Jun 13, 2023
Kind
B2
Abstract

A system and method including determining, for a specified target measure column of a first dataset including a plurality of records, the metadata of the first dataset, including a probability distribution for the specified target column and dimension scores for the dimensions for the first dataset conditioned on the specified target measure column, where the first dataset comprises a plurality of columns including the at least one target measure column and a plurality of non-numeric, dimension columns for the records of the first dataset; determining, for a subset of data of the first dataset based on one or more specified variables, dimension scores for the dimensions of the subset of data approximately derived from the determined metadata of the first dataset; and providing recommendations of top contributors based on the approximated dimension scores of dimensions of the subset of data.

Claims (68)

1. A computer-implemented method, the method comprising:

receiving all numeric values of a specified target measure column of a first dataset including a plurality of records, the first dataset having a plurality of columns including the specified target measure column and a plurality of non-numeric, dimension columns for the records of the first dataset;

discretizing each of the received numeric values of the target specified measure column into a plurality of bins, the plurality of bins being a pre-defined value, each of the bins having an equal interval width, and each of the bins having an index number;

generating a bin index column that contains a determined bin index number for each numeric value of the specified target measure column of each record in the first dataset;

determining a bin probability that represents a probability of each of the numeric values of the specified target measure column of the first dataset being in each of the bins based on the generated bin index column;

determining, based on the generated bin index column, a dimension score for each dimension column of the first dataset in each bin;

forming, based on the determined dimension score, a dimension score matrix for the first dataset; and

saving the determined bin probability and the dimension score matrix as metadata for the first dataset.

2. The method of claim 1 , further comprising:

receiving an indication of one or more specified variable values, the specified variable values each being a value selected from one or more of the plurality of non-numeric, dimension columns of the first dataset;

retrieving values in the bin index column of the records related to the specified one or more variable values;

determining a second bin probability, based on the retrieved bin index column values of the records related to the specified one or more variable values, that represents a probability of a value in the retrieved bin index column of the records related to the specified one or more variable values being in each of the bins;

deriving, by a first calculation, an approximated dimension score vector of a subset of the first dataset related to the specified one or more variable values based on the determined bin probability of a value being in each of the bins for the first dataset, the determined dimension score matrix for the first dataset, and the determined second bin probability; and

saving an output of the approximated dimension score vector.

3. The method of claim 2 , wherein a weighted dimension score vector is used in deriving the approximated dimension score vector.

4. The method of claim 2 , further comprising:

determining a first set of dimensions in the approximated dimension score vector having a highest value relative to each other, the number of dimension in the set being predefined; and

presenting the determined set of dimensions to a user.

5. The method of claim 4 , further comprising:

determining a second set of dimensions in the approximated dimension score vector having a highest value, the number of dimension in the second set being predefined and fewer than the number of dimensions in the first set; and

presenting the determined second set of dimensions to a user.

6. The method of claim 2 , wherein the first calculation to derive the approximated dimension score vector is substituted with a second calculation based on at least the determined bin probability for the first dataset and the determined dimension score vector for the first dataset.

7. A non-transitory, computer readable medium having executable instructions stored therein that, when executed by a computer processor cause the processor to perform a method, the method comprising:

receiving all numeric values of a specified target measure column of a first dataset including a plurality of records, the first dataset having a plurality of columns including the target measure column and a plurality of non-numeric, dimension columns for the records of the first dataset;

discretizing each of the received numeric values of the target specified measure column into a plurality of bins, the plurality of bins being a pre-defined value, each of the bins having an equal interval width, and each of the bins having an index number;

generating a bin index column that contains a determined bin index number for each numeric value of the specified target measure column of each record in the first dataset;

determining a bin probability that represents a probability of each of the numeric values of the specified target measure column of the first dataset in each of the bins based on the generated bin index column;

determining, based on the generated bin index column, a dimension score for each dimension column of the first dataset in each bin;

forming, based on the determined dimension score, a dimension score matrix for the first dataset; and

saving the determined bin probability and the dimension score matrix as metadata for the first dataset.

8. The medium of claim 7 , further comprising:

receiving an indication of one or more specified variable values, the specified variable values each being a value selected from one or more of the plurality of non-numeric, dimension columns of the first dataset;

retrieving values in the bin index column of the records related to the specified one or more variable values;

determining a second bin probability, based on the retrieved bin index column values of the records related to the specified one or more variable values, that represents a probability of a value in the retrieved bin index column of the records related to the specified one or more variable values being in each of the bins;

deriving, by a first calculation, an approximated dimension score vector of the first dataset related to the specified one or more variable values based on the determined bin probability of a value being in each of the bins for the first dataset, the determined dimension score matrix for the first dataset, and the determined second bin probability; and

saving an output of the approximated dimension score vector.

9. The medium of claim 8 , wherein a weighted dimension score vector is used in deriving the approximated dimension score vector.

10. The medium of claim 8 , further comprising:

determining a first set of dimensions in the approximated dimension score vector having a highest relative value, the number of dimension in the first set being predefined; and

presenting the determined first set of dimensions to a user.

11. The medium of claim 10 , further comprising:

determining a second set of dimensions in the approximated dimension score vector having a highest value, the number of dimension in the second set being predefined and fewer than the number of dimensions in the first set; and

presenting the determined second set of dimensions to a user.

12. The medium of claim 8 , wherein the first calculation to derive the approximated dimension score vector is substituted with a second calculation based on at least the determined bin probability for the first dataset and the determined dimension score vector for the first dataset.

13. A system, the system comprising:

a computer processor, and

computer memory, coupled to the computer processor, storing instructions that, when executed by the computer processor cause the computer processor to:

receive all numeric values of a specified target measure column of a first dataset including a plurality of records, the first dataset having a plurality of columns including the target measure column and a plurality of non-numeric, dimension columns for the records of the first dataset;

discretize each of the received numeric values of the specified target measure column into a plurality of bins, the plurality of bins being a pre-defined value, each of the bins having an equal interval width, and each of the bins having an index number;

generate a bin index column that contains a determined bin index number for each numeric value of the specified target measure column of each record in the first dataset;

determine a bin probability that represents a probability of each of the numeric values of the specified target measure column of the first dataset being in each of the bins based on the generated bin index column;

determine, based on the generated bin index column, a dimension score for each dimension column of the first dataset in each bin;

form, based on the determined dimension score, a dimension score matrix for the first dataset; and

save the determined bin probability and the dimension score matrix as metadata for the first dataset.

14. The system of claim 13 , further comprising:

receiving an indication of one or more specified variable values, the specified variable values each being a value selected from one or more of the plurality of non-numeric, dimension columns of the first dataset;

retrieving values in the bin index column of the records related to the specified one or more variable values;

determining a second bin probability, based on the retrieved bin index column values of the records related to the specified one or more variable values, that represents a probability of a value in the retrieved bin index column of the records related to the specified one or more variable values being in each of the bins;

deriving, by a first calculation, an approximated dimension score vector of the first dataset related to the specified one or more variable values based on the determined bin probability of a value being in each of the bins for the first dataset, the determined dimension score matrix for the first dataset, and the determined second bin probability; and

saving an output of the approximated dimension score vector.

15. The system of claim 14 , wherein a weighted dimension score vector is used in deriving the approximated dimension score vector.

16. The system of claim 14 , further comprising:

determining a first set of dimensions in the approximated dimension score vector having a highest value relative to each other, the number of dimension in the first set being predefined; and

presenting the determined first set of dimensions to a user.

17. The system of claim 16 , further comprising:

determining a second set of dimensions in the approximated dimension score vector having a highest value, the number of dimension in the second set being predefined and fewer than the number of dimensions in the first set; and

presenting the determined second set of dimensions to a user.

18. The system of claim 14 , wherein the first calculation to derive the approximated dimension score vector is substituted with a second calculation based on at least the determined bin probability for the first dataset and the determined dimension score vector for the first dataset.

Assignments (2)
CHANGE OF NAME Recorded Jan 26, 2026
From: BUSINESS OBJECTS SOFTWARE LIMITED
To: SAP IRELAND LIMITED
Reel/Frame 074510/0354 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2021
From: WU, YING; KAUFMANN, MALTE CHRISTIAN; MCSHANE, ALAN; BANERJEE, ANIRBAN; MAGUIRE, GARETH
To: BUSINESS OBJECTS SOFTWARE LTD.
Reel/Frame 056342/0627 →
Continuity (1)
Related Publication 20220382729A1 · Dec 1, 2022