Automated contribution analysis for question answering
This disclosure describes techniques and architecture provide automated contribution analysis for “why question” style NLQ answering, e.g., “why is revenue down in North America Q1 2022.” In particular, the techniques described herein combine multiple signals together including, for example, frequency of use of combinations of dimensions in previous NLQs (warm-start), statistical information about columns (e.g., entropy), correlation/co-occurrence between pairs of dimension columns, and correlation between dimensions and dates. This information is used with a set of heuristics and rules to pick the best set of dimensions as contributing factors for a particular metric over a particular time period and present an automatic contribution analysis to the users to give them insights into their data.
1 . A computer-implemented method comprising:
providing a dataset arranged in tabular form comprising rows and columns, wherein each column has a name representing a dimension;
identifying a topic related to a description of the dataset based on user input;
based on the topic, triggering a contributing dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the contribution dimension recommendation workflow comprises:
analyzing, using a knowledge discovery in databases (KDD) application programming interface (API), one or more heuristics comprising scanning visuals displayed to users related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for a plurality of dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and the dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses;
based on the analyzing, scoring each dimension;
based on each score, ranking each dimension; and
based on ranking each dimension, recommending the one or more dimensions for use in answering why-type NLQs;
receiving, from a user, a why-type NLQ;
based on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;
based on the one or more dimensions and the metric, searching the dataset for values related to the metric; and
based on the searching, providing results to the user with respect to the why-type NLQ.
2 . The computer-implemented method of claim 1 , wherein the metric has an aggregation of values that is one of sum, average, or count.
3 . The computer-implemented method of claim 1 , further comprising:
excluding, from the one or more dimensions, dimensions having a cardinality that is greater than or equal to a first threshold of sample size from the dataset and that is less than a second threshold that is less than the first threshold.
4 . The computer-implemented method of claim 1 , further comprising:
determining strongly-correlated dimensions with a co-occurrence greater than or equal to a threshold with respect to another dimension;
based on the strongly-correlated dimensions, determining a set of dimensions with a cardinality that are within a cardinality range;
selecting one dimension of the set of dimensions with a highest score; and adding values from the one dimension to the results.
5 . The computer-implemented method of claim 1 , further comprising:
determining strongly-correlated dimensions with a co-occurrence greater than or equal to a threshold with respect to another dimension;
based on the strongly-correlated dimensions, determining a set of dimensions with a cardinality that are within a cardinality range;
selecting one dimension of the set of dimensions with a highest score; and adding values from the one dimension to the results.
6 . The computer-implemented method of claim 1 , wherein the dimensions are columns of tables.
7 . A computer-implemented method comprising:
based at least in part on a topic related to a dataset, triggering a contribution dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the contribution dimension recommendation workflow comprises:
analyzing one or more heuristics;
ranking individual dimensions of the dataset; and
based at least in part on ranking the individual dimensions, recommending the one or more dimensions for use in answering why-type NLQs;
receiving, from a user, a why-type NLQ;
based at least in part on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;
based at least in part on the one or more dimensions and the metric, searching the dataset for one or more values related to the metric; and
based at least in part on the searching, providing results to the user with respect to the why-type NLQ.
8 . The computer-implemented method of claim 7 , wherein the metric has an aggregation of values that is one of sum, average, or count.
9 . The computer-implemented method of claim 7 , further comprising:
if a dimension of the one or more dimensions includes a filter, eliminating the dimension from the one or more dimensions.
10 . The computer-implemented method of claim 7 , further comprising:
excluding, from the one or more dimensions, dimensions with very high cardinality more than or equal to 95% of sample size from the dataset and very low (0, 1) cardinality.
11 . The computer-implemented method of claim 7 , wherein the heuristics comprise one or more heuristics comprising scanning visuals displayed to users related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for individual dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses.
12 . The computer-implemented method of claim 11 , wherein the contributing dimension recommendation workflow for the contribution analysis is performed by a knowledge discovery in databases (KDD) application programming interface (API).
13 . The computer-implemented method of claim 7 , wherein at least part of the contributing dimension recommendation workflow for the contribution analysis occurs during creation of the topic.
14 . The computer-implemented method of claim 13 , wherein at least part of the contributing dimension recommendation workflow for the contribution analysis occurs offline during creation of the topic.
15 . One or more computer-readable media storing computer-executable instructions that, when executed, cause one or more processors to perform operations comprising:
based at least in part on a topic related to a dataset, triggering a contributing dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the dimension recommendation workflow comprises:
analyzing one or more heuristics;
ranking individual dimensions of the dataset; and
based at least in part on ranking the individual dimensions, recommending the one or more dimensions for use in answering why-type NLQs;
receiving, from a user, a why-type NLQ;
based at least in part on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;
based at least in part on the one or more dimensions and the metric, searching the dataset for one or more values related to the metric; and
based at least in part on the searching, providing results to the user with respect to the why-type NLQ.
16 . The one or more computer-readable media of claim 15 , wherein the metric has an aggregation of values that is one of sum, average, or count.
17 . The one or more computer-readable media of claim 15 , wherein the operations further comprise:
if a dimension of the one or more dimensions includes a filter, eliminating the dimension from the one or more dimensions.
18 . The one or more computer-readable media of claim 15 , wherein the heuristics comprise one or more heuristics comprising scanning visuals displayed to user related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for the individual dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses.
19 . The one or more computer-readable media of claim 18 , wherein the contributing dimension recommendation workflow for the contribution analysis is performed by a knowledge discovery in databases (KDD) application programming interface (API).
20 . The one or more computer-readable media of claim 19 , wherein:
at least part of the contributing dimension recommendation workflow for the contribution analysis occurs during creation of the topic; and
at least part of the contributing dimension recommendation workflow for the contribution analysis occurs offline during creation of the topic.