IP Library › Granted Patent US 12,411,857
Granted Patent B2
US 12,411,857 · App. 18/164,992 · Granted Sep 9, 2025

Recommending aggregate questions in a conversational data exploration

Inventors: Ritwik Chaudhuri (Bangalore, IN); Rajmohan Chandrahasan (Perunagar, IN); Kirushikesh DB (Salem, IN); Arvind Agarwal (Delhi, IN)
Assignee: International Business Machines Corporation
G06F16/24578G06F16/221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,857
App. No.
18/164,992
Granted
Sep 9, 2025
Kind
B2
Abstract

Embodiments of the present invention provide an approach for exploring interesting data patterns in structured tables through recommending aggregate questions in a conversational data exploration. Specially, interesting features and operators are selected that are used to frame aggregate questions based on user intent and the data. The aggregate questions are ranked based on user persona and interestingness of the questions. The approach dynamically adapts and improves the recommendation of interesting and relevant aggregate questions for the user based on user feedback iteratively.

Claims (59)

1. A method for retrieving data in a conversational data exploration, comprising:

(a) calculating, by a processor, an importance score for each column within a dataset;

selecting, by the processor, a set of the columns based on the importance score of each column;

(b) generating, by the processor, a list of operators, wherein generating the list of operators includes:

determining Fisher scores for each column, and designating columns having a Fisher score above a threshold as part of a subset of data,

for the subset of data, calculating entropy values for a plurality of data slices of the subset of data, and

generating, for each column within the set of columns, the list of operators based on the Fisher scores and the calculated entropy values;

(c) generating, by the processor, a set of natural language questions based on the set of columns, the list of operators, and metadata related to the dataset;

(d) receiving a set of user selected questions from the set of natural language questions;

(e) calculating, by the processor, a ranking score for each question within the set of natural language questions using an entropy-based scoring method and based on the set of user selected questions;

(f) ranking, by the processor, the set of natural language questions based on each ranking score;

(g) presenting to the user, by the processor, a set of relevant questions from the set of natural language questions based on a predefined threshold;

(h) receiving another set of user selected questions from the set of relevant questions; and

(i) repeating (e) through (g) at least once.

2. The method of claim 1 , wherein the importance score of each column is calculated based on a role, responsibility, or intent of a user.

3. The method of claim 1 , wherein the metadata includes a set of column headers, a description, and a title related to the dataset.

4. The method of claim 1 , further comprising receiving, by the processor, a set of column headers from a user related to the dataset.

5. The method of claim 4 , further comprising presenting, by the processor, a subset of questions from the set of relevant questions based on the set of column headers.

6. The method of claim 1 , wherein the list of operators includes at least one of average, minimum, maximum, more than, less than, above, below, top K percent, fraction, total, majority, minority, missing, outlier, after, before, and within.

7. A computing system for retrieving data in a conversational data exploration, comprising:

a processor;

a memory device coupled to the processor; and

a computer readable storage device coupled to the processor, wherein the storage device contains program code executable by the processor via the memory device to implement a method, the method comprising:

(a) calculating, by a processor, an importance score for each column within a dataset;

selecting, by the processor, a set of the columns based on the importance score of each column;

(b) generating, by the processor, a list of operators, wherein generating the list of operators includes:

determining Fisher scores for each column, and designating columns having a Fisher score above a threshold as part of a subset of data,

for the subset of data, calculating entropy values for a plurality of data slices of the subset of data, and

generating, for each column within the set of columns, the list of operators based on the Fisher scores and the calculated entropy values;

(c) generating, by the processor, a set of natural language questions based on the set of columns, the list of operators, and metadata related to the dataset;

(d) receiving a set of user selected questions from the set of natural language questions;

(e) calculating, by the processor, a ranking score for each question within the set of natural language questions using an entropy-based scoring method and based on the set of user selected questions;

(f) ranking, by the processor, the set of natural language questions based on each ranking score;

(g) presenting to the user, by the processor, a set of relevant questions from the set of natural language questions based on a predefined threshold;

(h) receiving another set of user selected questions from the set of relevant questions; and

(i) repeating (e) through (g) at least once.

8. The computing system of claim 7 , wherein the importance score of each column is calculated based on a role, responsibility, or intent of a user.

9. The computing system of claim 7 , wherein the metadata includes a set of column headers, a description, and a title related to the dataset.

10. The computing system of claim 7 , further comprising receiving, by the processor, a set of column headers from a user related to the dataset.

11. The computing system of claim 10 , further comprising presenting, by the processor, a subset of questions from the set of relevant questions based on the set of column headers.

12. The computing system of claim 7 , wherein the list of operators includes at least one of average, minimum, maximum, more than, less than, above, below, top K percent, fraction, total, majority, minority, missing, outlier, after, before, and within.

13. A computer program product for retrieving data in a conversational data exploration, the computer program product comprising a computer readable storage device, and program instructions stored on the computer readable storage device, to:

(a) calculate, by a processor, an importance score for each column within a dataset;

selecting, by the processor, a set of the columns based on the importance score of each column;

(b) generate, by the processor, a list of operators, wherein generating the list of operators includes:

determining Fisher scores for each column, and designating columns having a Fisher score above a threshold as part of a subset of data,

for the subset of data, calculating entropy values for a plurality of data slices of the subset of data, and

generating, for each column within the set of columns, the list of operators based on the Fisher scores and the calculated entropy values;

(c) generate, by the processor, a set of natural language questions based on the set of columns, the list of operators, and metadata related to the dataset;

(d) receive a set of user selected questions from the set of natural language questions;

(e) calculate, by the processor, a ranking score for each question within the set of natural language questions using an entropy-based scoring method and based on the set of user selected questions;

(f) rank, by the processor, the set of natural language questions based on each ranking score;

(g) present to the user, by the processor, a set of relevant questions from the set of natural language questions based on a predefined threshold;

(h) receive another set of user selected questions from the set of relevant questions; and

(i) repeat (e) through (g) at least once.

14. The computer program product of claim 13 , wherein the importance score of each column is calculated based on a role, responsibility, or intent of a user.

15. The computer program product of claim 13 , wherein the metadata includes a set of column headers, a description, and a title related to the dataset.

16. The computer program product of claim 13 , further comprising program instructions stored on the computer readable storage device to receive, by the processor, a set of column headers from a user related to the dataset.

17. The computer program product of claim 16 , further comprising program instructions stored on the computer readable storage device to present, by the processor, a subset of questions from the set of relevant questions based on the set of column headers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2023
From: CHAUDHURI, RITWIK; CHANDRAHASAN, RAJMOHAN; DB, KIRUSHIKESH; AGARWAL, ARVIND
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 062603/0380 →
Continuity (1)
Related Publication 20240265020A1 · Aug 8, 2024
References Cited (22)
US 9971967B2 · Bufe, III et al. · 2018 [cited by applicant]
US 11847424B1 · Harkous · 2023 [cited by examiner]
US 20140075334A1 · Dror · 2014 [cited by examiner]
US 20150169544A1 · Bufe, III et al. · 2015 [cited by applicant]
US 20150371548A1 · Samid · 2015 [cited by examiner]
US 20170235848A1 · Van Dusen · 2017 [cited by examiner]
US 20190164063A1 · Moura · 2019 [cited by examiner]
US 20190272296A1 · Prakash · 2019 [cited by examiner]
US 20200134019A1 · Podgorny · 2020 [cited by examiner]
US 20200159772A1 · Zoumpoulakis · 2020 [cited by examiner]
US 20210019309A1 · Yadav · 2021 [cited by examiner]
US 20210374168A1 · Srinivasan · 2021 [cited by examiner]
Authors et al.: Disclosed Anonymously, “Method for Generating, Selecting, and Ranking Natural Language Modifiers as Starting Points for Datasets in Business Intelligence Systems”, IP.com No. IPCOM000264234D, IP.com Elec… [cited by applicant]
Arjun Srinivasan et al., “Snowy: Recommending Utterances for Conversational Visual Analysis”, Publication Date Oct. 12, 2021, pp. 864-880. [cited by applicant]
Saichandra Pandraju et al., “Answer-Aware Question Generation from Tabular and Textual Data using T5”, Publication Date Sep. 20, 2019, pp. 256-267. [cited by applicant]
Kedar Dhamdhere et al., “Analyza: Exploring Data with Conversation”, Publication Date Mar. 7, 2017, pp. 493-504. [cited by applicant]
Zhihui Yang et al., “iExplore: Accelerating Exploratory Data Analysis by Predicting User Intention”, Publication Date May 12, 2018, pp. 149-165. [cited by applicant]
Milo, “Automating Exploratory Data Analysis via Machine Learning: An Overview”, SIGMOD '20, Jun. 14-19, 2020, Portland, OR, USA, 6 pgs. [cited by applicant]
Milo, “Deep Reinforcement-LearningFramework for Exploratory Data Analysis”, aiDM'18, Jun. 10, 2018, Houston, TX, USA, 4 pgs. [cited by applicant]
Vartak et al., “Efficient Data-Driven Visualization Recommendations to Support Visual Analytics”, Proceedings VLDB Endowment. Sep. 2015 ; 8(13): 2182-2193, 41 pgs. [cited by applicant]
Singh, Exploratory Data Analysis with Tableau, Proceedings VLDB Endowment. Sep. 2015 ; Proceedings VLDB Endowment. Sep. 2015, www.pluralsight.com/guides/exploratory-data-analysis-with-tableau, Jun. 24, 2020, 12 pgs. [cited by applicant]
Microsoft, “Turn your data into immediate impact”, https://powerbi.microsoft.com/en-au/, Oct. 2022, 12 pgs. [cited by applicant]