IP Library › Granted Patent US 10,706,045
Granted Patent B1
US 10,706,045 · App. 16/387,016 · Granted Jul 7, 2020

Natural language querying of a data lake using contextualized knowledge bases

Inventors: Kanav Hasija (Nagar New Delhi, IN); Kartik R. Sayani (Ghaziabad, IN)
Assignee: INNOVACCER INC.
G06F16/243G06F16/245G06F16/248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,706,045
App. No.
16/387,016
Filed
Apr 17, 2019
Granted
Jul 7, 2020
Kind
B1
Art Unit
2167
USPC
707/722
Abstract

A method of querying a data lake using natural language includes: receiving a natural language query directed to an electronic data lake; parsing the natural language query to determine a plurality of entities within the natural language query; identifying the plurality of entities using at least one contextual knowledge base, wherein the plurality of entities are compared against at least one entry in the at least one contextual knowledge base; mapping a dependency of the plurality of identified entities based on the parsed natural language query; constructing a structured data query based on the plurality of identified entities and the mapped dependency; and automatically generating a visual output of a result of the structured data query based on at least one characteristic from the set of: a data type, a data format, and a data size of the result of the structured data query.

Claims (39)

1. A method of querying a data lake using natural language, comprising the following steps:

receiving a natural language query directed to an electronic data lake;

parsing the natural language query to determine a plurality of entities within the natural language query;

identifying the plurality of entities using at least one contextual knowledge base, wherein the plurality of entities are tabulated in at least one of a plurality of data tables by entity type and compared against at least one entry in the at least one contextual knowledge base, wherein at least one phrase of the natural language query is combined, the plurality of entities are soft-matched, and at least one entity above a threshold confidence level is identified, and wherein a relationship table structure knowledge base provides a relationship between at least two of the plurality of entities by determining relational links at least two of the plurality of data tables;

mapping a dependency relationship between the plurality of identified entities to determine relational parts of speech of the plurality of identified entities based on the parsed natural language query;

constructing a structured data query based on the plurality of identified entities and the mapped dependency; and

automatically generating a visual output of a result of the structured data query, wherein a format of the visual output is recommended by a visual recommender knowledge base based on at least: a number of columns or a number of rows of the result of the structured data query.

2. The method of claim 1 , wherein the at least one contextual knowledge base is selected from the set of: entity, entity thesaurus, code set, and aggregation and slice identifier knowledge bases.

3. The method of claim 1 , wherein the visual output is a type selected from the set of: pie graph, bar graph, line graph, X-Y plot, area chart, scatter plot, bubble plot, and histogram.

4. The method of claim 3 , comprising at least two types of visual output.

5. The method of claim 1 , wherein the result of the structured data is saved as a file.

6. The method of claim 1 , wherein the step of parsing the natural language query comprises the following steps:

parsing at least one sentence of the natural language query;

parsing at least one part of speech of the at least one sentence; and

creating a dependency tree based on the at least one part of speech.

7. The method of claim 1 , further comprising the step of suggesting at least a portion of the natural language query.

8. The method of claim 1 , wherein the structured data query is constructed in Structured Query Language (SQL).

9. A computerized system having a memory and a processor for querying a data lake using natural language, comprising:

an electronic data lake;

at least one contextual knowledge base; and

a user computer device in communication with the electronic data lake and the at least one contextual knowledge base over at least one network, wherein the user computer device has a processor and a non-transitory memory, wherein the processor of the user computer device is configured to:

receive a natural language query directed to the electronic data lake;

parse the natural language query to determine a plurality of entities within the natural language query;

identify the plurality of entities using the at least one contextual knowledge base, wherein the plurality of entities are tabulated in at least one of a plurality of data tables by entity type and compared against at least one entry in the at least one contextual knowledge base, wherein at least one phrase of the natural language query is combined, the plurality of entities are soft-matched, and at least one entity above a threshold confidence level is identified, and wherein a relationship table structure knowledge base provides a relationship between at least two of the plurality of entities by determining relational links at least two of the plurality of data tables;

map a dependency relationship between the plurality of identified entities to determine relational parts of speech of the plurality of identified entities based on the parsed natural language query;

construct a structured data query based on the plurality of identified entities and the mapped dependency; and

automatically generate a visual output of a result of the structured data query on a display of the user computer device, wherein the visual output is recommended by a visual recommender knowledge base based on at least: a number of columns or a number of rows of the result of the structured data query.

10. The system of claim 9 , wherein the at least one contextual knowledge base is selected from the set of: entity, entity thesaurus, code set, and aggregation and slice identifier knowledge bases.

11. The system of claim 9 , wherein the visual output is a type selected from the set of: pie graph, bar graph, line graph, X-Y plot, area chart, scatter plot, bubble plot, and histogram.

12. The system of claim 11 , comprising at least two types of visual output.

13. The system of claim 9 , wherein the result of the structured data is saved as a file.

14. The system of claim 9 , wherein parsing the natural language query comprises:

parsing at least one sentence of the natural language query;

parsing at least one part of speech of the at least one sentence; and

creating a dependency tree based on the at least one part of speech.

15. The system of claim 9 , wherein the user computer device is configured to suggest at least a portion of the natural language query.

16. The system of claim 9 , wherein the structured data query is constructed in Structured Query Language (SQL).

17. The system of claim 9 , wherein the display is selected from the set of: monitor, projector, screen, and wearable.

18. The system of claim 9 , wherein the user computer device is configured to return a tabular data output displayed with the visual output.

Assignments (2)
SECURITY INTEREST Recorded Apr 30, 2024
From: INNOVACCER INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 067267/0637 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2019
From: HASIJA, KANAV; SAYANI, KARTIK R.
To: INNOVACCER INC.
Reel/Frame 049056/0575 →
Priority Claims (1)
IN 201921005332 · Feb 11, 2019 · national
Cited By (14)
US 12,288,039 US 12,299,022 US 12,423,525 US 12,462,114 US 12,468,694 US 12,505,093 US 12,608,416 US 12,614,042 US 12,632,445 US 12,664,160 US 12,675,475 US 12,681,947 US 12,681,997 US 12,717,828