IP Library › Granted Patent US 12,626,158
Granted Patent B1
US 12,626,158 · App. 18/070,147 · Granted May 12, 2026

Automated contribution analysis for question answering

Inventors: Wojciech Aleksander Wilk (Austin, TX); Shannon Kalisky (Dripping Springs, TX); Rishav Chakravarti (White Plains, NY); William Michael Siler (Germantown, TN); Stephen Michael Ash (Seattle, WA); Rajesh Patel (Austin, TX); Joshua Noah Malters (Seattle, WA); Gregory David Adams (Seattle, WA); Jose Kunnackal John (Bellevue, WA)
Assignee: Amazon Technologies, Inc.
G06N5/04G06F40/295G06F40/30G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,158
App. No.
18/070,147
Granted
May 12, 2026
Kind
B1
Abstract

This disclosure describes techniques and architecture provide automated contribution analysis for “why question” style NLQ answering, e.g., “why is revenue down in North America Q1 2022.” In particular, the techniques described herein combine multiple signals together including, for example, frequency of use of combinations of dimensions in previous NLQs (warm-start), statistical information about columns (e.g., entropy), correlation/co-occurrence between pairs of dimension columns, and correlation between dimensions and dates. This information is used with a set of heuristics and rules to pick the best set of dimensions as contributing factors for a particular metric over a particular time period and present an automatic contribution analysis to the users to give them insights into their data.

Claims (59)

1 . A computer-implemented method comprising:

providing a dataset arranged in tabular form comprising rows and columns, wherein each column has a name representing a dimension;

identifying a topic related to a description of the dataset based on user input;

based on the topic, triggering a contributing dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the contribution dimension recommendation workflow comprises:

analyzing, using a knowledge discovery in databases (KDD) application programming interface (API), one or more heuristics comprising scanning visuals displayed to users related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for a plurality of dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and the dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses;

based on the analyzing, scoring each dimension;

based on each score, ranking each dimension; and

based on ranking each dimension, recommending the one or more dimensions for use in answering why-type NLQs;

receiving, from a user, a why-type NLQ;

based on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;

based on the one or more dimensions and the metric, searching the dataset for values related to the metric; and

based on the searching, providing results to the user with respect to the why-type NLQ.

2 . The computer-implemented method of claim 1 , wherein the metric has an aggregation of values that is one of sum, average, or count.

3 . The computer-implemented method of claim 1 , further comprising:

excluding, from the one or more dimensions, dimensions having a cardinality that is greater than or equal to a first threshold of sample size from the dataset and that is less than a second threshold that is less than the first threshold.

4 . The computer-implemented method of claim 1 , further comprising:

determining strongly-correlated dimensions with a co-occurrence greater than or equal to a threshold with respect to another dimension;

based on the strongly-correlated dimensions, determining a set of dimensions with a cardinality that are within a cardinality range;

selecting one dimension of the set of dimensions with a highest score; and adding values from the one dimension to the results.

5 . The computer-implemented method of claim 1 , further comprising:

determining strongly-correlated dimensions with a co-occurrence greater than or equal to a threshold with respect to another dimension;

based on the strongly-correlated dimensions, determining a set of dimensions with a cardinality that are within a cardinality range;

selecting one dimension of the set of dimensions with a highest score; and adding values from the one dimension to the results.

6 . The computer-implemented method of claim 1 , wherein the dimensions are columns of tables.

7 . A computer-implemented method comprising:

based at least in part on a topic related to a dataset, triggering a contribution dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the contribution dimension recommendation workflow comprises:

analyzing one or more heuristics;

ranking individual dimensions of the dataset; and

based at least in part on ranking the individual dimensions, recommending the one or more dimensions for use in answering why-type NLQs;

receiving, from a user, a why-type NLQ;

based at least in part on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;

based at least in part on the one or more dimensions and the metric, searching the dataset for one or more values related to the metric; and

based at least in part on the searching, providing results to the user with respect to the why-type NLQ.

8 . The computer-implemented method of claim 7 , wherein the metric has an aggregation of values that is one of sum, average, or count.

9 . The computer-implemented method of claim 7 , further comprising:

if a dimension of the one or more dimensions includes a filter, eliminating the dimension from the one or more dimensions.

10 . The computer-implemented method of claim 7 , further comprising:

excluding, from the one or more dimensions, dimensions with very high cardinality more than or equal to 95% of sample size from the dataset and very low (0, 1) cardinality.

11 . The computer-implemented method of claim 7 , wherein the heuristics comprise one or more heuristics comprising scanning visuals displayed to users related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for individual dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses.

12 . The computer-implemented method of claim 11 , wherein the contributing dimension recommendation workflow for the contribution analysis is performed by a knowledge discovery in databases (KDD) application programming interface (API).

13 . The computer-implemented method of claim 7 , wherein at least part of the contributing dimension recommendation workflow for the contribution analysis occurs during creation of the topic.

14 . The computer-implemented method of claim 13 , wherein at least part of the contributing dimension recommendation workflow for the contribution analysis occurs offline during creation of the topic.

15 . One or more computer-readable media storing computer-executable instructions that, when executed, cause one or more processors to perform operations comprising:

based at least in part on a topic related to a dataset, triggering a contributing dimension recommendation workflow for a contribution analysis to determine ranks of one or more dimensions of the dataset that contribute to answering why-type natural language questions (NLQs), wherein the dimension recommendation workflow comprises:

analyzing one or more heuristics;

ranking individual dimensions of the dataset; and

based at least in part on ranking the individual dimensions, recommending the one or more dimensions for use in answering why-type NLQs;

receiving, from a user, a why-type NLQ;

based at least in part on intent representation (IR) with respect to the why-type NLQ, selecting a metric related to the why-type NLQ;

based at least in part on the one or more dimensions and the metric, searching the dataset for one or more values related to the metric; and

based at least in part on the searching, providing results to the user with respect to the why-type NLQ.

16 . The one or more computer-readable media of claim 15 , wherein the metric has an aggregation of values that is one of sum, average, or count.

17 . The one or more computer-readable media of claim 15 , wherein the operations further comprise:

if a dimension of the one or more dimensions includes a filter, eliminating the dimension from the one or more dimensions.

18 . The one or more computer-readable media of claim 15 , wherein the heuristics comprise one or more heuristics comprising scanning visuals displayed to user related to the dataset, scanning previously asked questions, computing entropy of a distribution of values for the individual dimensions of the dataset, previously run anomaly detections, co-occurrence between measures and dimensions using logistic regression, and frequency of use of the dimensions in previous contribution analyses.

19 . The one or more computer-readable media of claim 18 , wherein the contributing dimension recommendation workflow for the contribution analysis is performed by a knowledge discovery in databases (KDD) application programming interface (API).

20 . The one or more computer-readable media of claim 19 , wherein:

at least part of the contributing dimension recommendation workflow for the contribution analysis occurs during creation of the topic; and

at least part of the contributing dimension recommendation workflow for the contribution analysis occurs offline during creation of the topic.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2026
From: WILK, WOJCIECH ALEKSANDER; KALISKY, SHANNON; CHAKRAVARTI, RISHAV; SILER, WILLIAM MICHAEL; ASH, STEPHEN MICHAEL; PATEL, RAJESH; ADAMS, GREGORY DAVID; JOHN, JOSE KUNNACKAL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 074271/0098 →
References Cited (25)
US 9159034B2 · Pinckney · 2015 [cited by examiner]
US 12001467B1 · U et al. · 2024 [cited by applicant]
US 12223080B1 · Al-Rikabi · 2025 [cited by examiner]
US 20080189655A1 · Kol et al. · 2008 [cited by applicant]
US 20140101154A1 · Nagulakonda et al. · 2014 [cited by applicant]
US 20150269176A1 · Marantz et al. · 2015 [cited by applicant]
US 20150339590A1 · Maarek · 2015 [cited by examiner]
US 20190272296A1 · Prakash · 2019 [cited by examiner]
US 20200012892A1 · Goodsitt et al. · 2020 [cited by applicant]
US 20210019309A1 · Yadav et al. · 2021 [cited by applicant]
US 20210374123A1 · Nicol et al. · 2021 [cited by applicant]
US 20220035943A1 · Jones · 2022 [cited by applicant]
US 20220067281A1 · Hu et al. · 2022 [cited by applicant]
US 20220342583A1 · Scott et al. · 2022 [cited by applicant]
US 20230070715A1 · Pajak et al. · 2023 [cited by applicant]
US 20230274098A1 · Syeda-Mahmood · 2023 [cited by applicant]
US 20230394031A1 · Kothari et al. · 2023 [cited by applicant]
US 20230418873A1 · Gao et al. · 2023 [cited by applicant]
Office Action for U.S. Appl. No. 18/070,117, mailed on Nov. 27, 2024, 20 pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/070,170, Dated Oct. 30, 2024, Ash, “Synthetic Question Generation for Use With Natural Language Question Searching,” 30 pages. [cited by applicant]
Fernandez, Raul Castro, et al. “Seeping semantics: Linking datasets using word embeddings for data discovery.” IEEE 34th International Conference on Data Engineering (ICDE) 2018, pp. 989-1000, Article 8509314. [cited by applicant]
Hulsebos, Madelon, et al. “Sherlock: A deep learning approach to semantic data type detection.” (Semantic typing as supervised learning) KDD '19: Proceedings of the 25t ACM SIGKDD International Conference on Knowledge D… [cited by applicant]
Petrovski et al., “Embedding Individual Table Columns for Resilient SQL Chatbots”, arXiv:1811.00633v1, Nov. 1, 2018, 7 pages. [cited by applicant]
Suhara, Yoshihiko, et al., “Annotating Columns with Pre-trained Language Models” SIGMOD '22: Proceeding of the 2022 International Conference on Management of Data, 2021, pp. 1493-1503. [cited by applicant]
Zhang, Dan, et al. “Sato: Contextual Semantic Type Detection in Tables.” (Follow-up paper to Sherlock) Proceeding of the VLDB Endowment, vol. 13, Issue 12, pp. 1835-1848, 2020, VLDB Endowment. [cited by applicant]