IP Library Granted Patent US 9,940,384
Granted Patent B2
US 9,940,384 · App. 14/969,211 · Granted Apr 10, 2018

Statistical clustering inferred from natural language to drive relevant analysis and conversation with users

Inventors: Stephen D. Gibson (Kemptville, CA); Alireza Pourshahid (Ottawa, CA); Vinay N. Wadhwa (Ottawa, CA); Graham A. Watts (Ottawa, CA)
Assignee: International Business Machines Corporation
G06F17/30598G06F17/28G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,940,384
App. No.
14/969,211
Granted
Apr 10, 2018
Kind
B2
Abstract

A mechanism is provided in a data processing system for statistical clustering inferred from natural language to drive relevant analysis. The mechanism receives a natural language text from a user and processes the natural language text to identify an entity of interest and a focus of statistical analysis. The mechanism performs a follow-up question and answer conversation with the user to receiving from the user one or more driving factor values for the one or more driving factors. The mechanism determines at least one cluster of entities matching the one or more driving factor values and generates at least one data visualization of the data in the corpus for the focus of statistical analysis having a scope that is narrowed based on the at least one cluster of entities matching the one or more driving factor values.

Claims (37)

1. A method, in a data processing system, for statistical clustering inferred from natural language to drive relevant analysis, the method comprising:

receiving a natural language text from a user;

processing the natural language text to identify an entity of interest and a focus of statistical analysis;

performing a follow-up question and answer conversation with the user to receive from the user one or more driving factor values for one or more driving factors for the focus of the statistical analysis;

determining at least one cluster of entities matching the one or more driving factor values; and

generating at least one data visualization of the data in a corpus for the focus of statistical analysis having a scope that is narrowed based on the at least one cluster of entities matching the one or more driving factor values.

2. The method of claim 1 , wherein performing a follow-up question and answer conversation comprises performing a clustering operation on data in the corpus for the focus of statistical analysis and determining one or more driving factors for the focus of the statistical analysis based on results of the clustering operation.

3. The method of claim 2 , wherein performing a follow-up question and answer conversation comprises detecting the most important driving factors for the focus of analysis based on the results of the clustering operation.

4. The method of claim 1 , wherein performing a follow-up question and answer conversation comprises generating one or more follow-up questions to be presented to the user to gather information required about the entity of interest and receiving responses to the one or more questions from the user.

5. The method of claim 4 , wherein performing a follow-up question and answer conversation further comprises parsing the responses to determine values for attributes that form the driving factors.

6. The method of claim 4 , wherein generating one or more follow-up questions comprises using slot tiller templates to generate the follow-up questions.

7. The method of claim 1 , wherein determining at least one cluster of entities matching the one or more driving factor values comprises creating clusters based on the driving factors and matching the entity of interest to at least one of the clusters.

8. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a computing device, causes the computing device to:

receive a natural language text from a user;

process the natural language text to identify an entity of interest and a focus of statistical analysis;

perform a follow-up question and answer conversation with the user to receive from the user one or more driving factor values for one or more driving factors for the focus of the statistical analysis;

determine at least one cluster of entities matching the one or more driving factor values; and

generate at least one data visualization of the data in a corpus for the focus of statistical analysis having a scope that is narrowed based on the at least one cluster of entities matching the one or more driving factor values.

9. The computer program product of claim 8 , wherein performing a follow-up question and answer conversation comprises performing a clustering operation on data in the corpus for the focus of statistical analysis and determining one or more driving factors for the focus of the statistical analysis based on results of the clustering operation.

10. The computer program product of claim 9 , wherein performing a follow-up question and answer conversation comprises detecting the most important driving factors for the focus of analysis based on the results of the clustering operation.

11. The computer program product of claim 8 , wherein performing a follow-up question and answer conversation comprises generating one or more follow-up questions to be presented to the user to gather information required about the entity of interest and receiving responses to the one or more questions from the user.

12. The computer program product of claim 11 , wherein performing a follow-up question and answer conversation further comprises parsing the responses to determine values for attributes that form the driving factors.

13. The computer program product of claim 11 , wherein generating one or more follow-up questions comprises using slot filler templates to generate the follow-up questions.

14. The computer program product of claim 8 , wherein determining at least one cluster of entities matching the one or more driving factor values comprises creating clusters based on the driving factors and matching the entity of interest to at least one of the clusters.

15. An apparatus comprising:

a processor; and

a memory coupled to the processor, wherein the memory comprises instructions which, When executed by the processor, cause the processor to:

receive a natural language text from a user;

process the natural language text to identify an entity of interest and a focus of statistical analysis;

perform a follow-up question and answer conversation with the user to receive from the user one or more driving factor values for one or more driving factors for the focus of the statistical analysis;

determine at least one cluster of entities matching the one or more driving factor values; and

generate at least one data visualization of the data in a corpus for the focus of statistical analysis having a scope that is narrowed based on the at least one cluster of entities matching the one or more driving factor values.

16. The apparatus of claim 15 , wherein performing a follow-up question and answer conversation comprises performing a clustering operation on data in the corpus for the focus of statistical analysis and determining one or more driving factors for the focus of the statistical analysis based on results of the clustering operation.

17. The apparatus of claim 16 , wherein performing a follow-up question and answer conversation comprises detecting the most important driving factors for the focus of analysis based on the results of the clustering operation.

18. The apparatus of claim 15 , wherein performing a follow-up question and answer conversation comprises generating one or more follow-up questions to be presented to the user to gather information required about the entity of interest and receiving responses to the one or more questions from the user.

19. The apparatus of claim 18 , wherein performing a follow-up question and answer conversation further comprises parsing the responses to determine values for attributes that form the driving factors.

20. The apparatus of claim 15 , wherein determining at least one cluster of entities matching the one or more driving factor values comprises creating clusters based on the driving factors and matching the entity of interest to at least one of the clusters.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.
Reel/Frame 058711/0689 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2015
From: GIBSON, STEPHEN D.; POURSHAHID, ALIREZA; WADHWA, VINAY N.; WATTS, GRAHAM A.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 037292/0768 →
Continuity (1)
Related Publication 20170169094A1 · Jun 15, 2017