IP Library › Granted Patent US 11,734,328
Granted Patent B2
US 11,734,328 · App. 16/195,471 · Granted Aug 22, 2023

Artificial intelligence based corpus enrichment for knowledge population and query response

Inventors: Chinnappa Guggilla (Bangalore, IN); Praneeth Shishtla (Bangalore, IN); Madhura Shivaram (Bangalore, IN)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06F16/353G06F16/288G06F16/3347G06F16/93G06N5/02G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,734,328
App. No.
16/195,471
Granted
Aug 22, 2023
Kind
B2
Abstract

In some examples, artificial intelligence based corpus enrichment for knowledge population and query response may include generating, based on annotated training documents, an entity and relation annotation model, identifying, based on application of the entity and relation annotation model to a document set that is to be annotated, entities and relations between the entities for each document of the document set to generate an annotated document set, and categorizing each annotated document into a plurality of categories. Artificial intelligence based corpus enrichment may include determining whether an identified category includes a specified number of annotated documents, and if not, additional annotated documents may be generated for the identified category that may represent a corpus. Further, artificial intelligence based corpus enrichment may include training, using the corpus, an artificial intelligence based decision support model, and utilizing the artificial intelligence based decision support model to respond to an inquiry.

Claims (112)

1. An artificial intelligence based corpus enrichment for knowledge population and query response apparatus comprising:

an entity and relation annotator, executed by at least one hardware processor, to

generate, based on annotated training documents, an entity and relation annotation model,

ascertain, a document set that is to be annotated, wherein documents of the document set include unstructured documents and semi-structured documents,

identify, based on application of the entity and relation annotation model to the document set, entities and relations between the entities for each document of the document set to generate an annotated document set,

determine, for each identified entity of the identified entities, an entity confidence score,

determine, for each identified relation of the identified relations between the entities, a relation confidence score,

identify, based on the entity confidence score, an entity that includes an entity confidence score that is less than an entity confidence score threshold,

identify, based on the relation confidence score, a relation that includes a relation confidence score that is less than a relation confidence score threshold,

generate another inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold, and

train, based on a response to the other inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold, the entity and relation annotation model;

a document categorizer, executed by the at least one hardware processor, to

categorize each annotated document of the annotated document set into a respective category of a plurality of categories;

a corpus generator and enricher, executed by the at least one hardware processor, to

identify a category of the plurality of categories,

determine whether the identified category includes a specified number of annotated documents, and

based on a determination that the identified category does not include the specified number of annotated documents, generate, for the identified category, additional annotated documents, wherein the annotated documents and the additional annotated documents of the identified category together represent a corpus;

an artificial intelligence model generator, executed by the at least one hardware processor, to

train, using the corpus, an artificial intelligence based decision support model; and

an inquiry response generator, executed by the at least one hardware processor, to

ascertain an inquiry related to an entity of the corpus, and

generate, by invoking the artificial intelligence based decision support model, a response to the inquiry.

2. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 1 , wherein the entity and relation annotator is further executed by the at least one hardware processor to generate, based on the annotated training documents, the entity and relation annotation model by:

transforming each annotated training document of the annotated training documents into a vector representation; and

generating, based on vector representations of the annotated training documents, the entity and relation annotation model.

3. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 1 , wherein the entity and relation annotator is further executed by the at least one hardware processor to:

identify entities for which a difference between entity confidence scores is less than a specified numerical value;

identify relations for which a difference between relation confidence scores is less than the specified numerical value;

generate another inquiry for verification of the entities for which the difference between the entity confidence scores is less than the specified numerical value, and the relations for which the difference between the relation confidence scores is less than the specified numerical value; and

train, based on a response to the other inquiry for verification of the entities for which the difference between the entity confidence scores is less than the specified numerical value, and the relations for which the difference between the relation confidence scores is less than the specified numerical value, the entity and relation annotation model.

4. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 1 , wherein the document categorizer is further executed by the at least one hardware processor to:

categorize each document of the document set that is to be annotated into the respective category of the plurality of categories,

wherein the entity and relation annotator is further executed by the at least one hardware processor to identify, based on application of the entity and relation annotation model to the document set, entities and relations between the entities for each document of the document set to generate the annotated document set by:

identifying, based on application of the entity and relation annotation model to documents of the identified category, entities and relations between the entities for each document of the identified category to generate annotated documents for the identified category.

5. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 1 , wherein the document categorizer is further executed by the at least one hardware processor to categorize each annotated document of the annotated document set into the respective category of the plurality of categories by:

transforming each annotated document of the annotated document set into an entity vector;

grouping, based on the entity vector for each annotated document, semantically similar entities; and

categorizing, based on the grouping, each annotated document of the annotated document set into the respective category of the plurality of categories.

6. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 1 , wherein the corpus generator and enricher is further executed by the at least one hardware processor to generate, for the identified category, additional annotated documents by:

segmenting the corpus into a plurality of sections to generate a preprocessed and segmented entity-annotated textual corpus.

7. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 6 , wherein for an invoice document, the plurality of sections include header, body, payment, and reference.

8. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 6 , wherein the corpus generator and enricher is further executed by the at least one hardware processor to generate, for the identified category, additional annotated documents by:

transforming each segment of the preprocessed and segmented entity-annotated textual corpus into a character representation to generate a plurality of character representations;

consolidating, based on the character representations, segments of the preprocessed and segmented entity-annotated textual corpus; and

learning, from each segment of the segments, character embeddings.

9. The artificial intelligence based corpus enrichment for knowledge population and query response apparatus according to claim 8 , wherein the corpus generator and enricher is further executed by the at least one hardware processor to generate, for the identified category, additional annotated documents by:

ascertaining a seed document;

segmenting the seed document into another plurality of sections;

transforming each segment of the segmented seed document into character represented vector embeddings;

generating a plurality of corresponding segments specific to each transformed segment of the segmented seed document; and

generating, based on the plurality of corresponding segments and the learned character embeddings, the additional annotated documents.

10. A computer implemented method for artificial intelligence based corpus enrichment for knowledge population and query response comprising:

generating, by an entity and relation annotator that is executed by at least one hardware processor, based on annotated training documents, an entity and relation annotation model;

ascertaining, by the entity and relation annotator that is executed by the at least one hardware processor, a document set that is to be annotated;

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, based on application of the entity and relation annotation model to the document set, entities and relations between the entities for each document of the document set to generate an annotated document set;

determining, by the entity and relation annotator that is executed by the at least one hardware processor, for each identified entity of the identified entities, an entity confidence score;

determining, by the entity and relation annotator that is executed by the at least one hardware processor, for each identified relation of the identified relations between the entities, a relation confidence score;

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, based on the entity confidence score, an entity that includes an entity confidence score that is less than an entity confidence score threshold;

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, based on the relation confidence score, a relation that includes a relation confidence score that is less than a relation confidence score threshold;

generating, by the entity and relation annotator that is executed by the at least one hardware processor, another inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold;

training, by the entity and relation annotator that is executed by the at least one hardware processor, based on a response to the other inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold, the entity and relation annotation model;

categorizing, by a document categorizer that is executed by the at least one hardware processor, each annotated document of the annotated document set into a respective category of a plurality of categories;

identifying, by a corpus generator and enricher that is executed by the at least one hardware processor, a category of the plurality of categories;

determining, by the corpus generator and enricher that is executed by the at least one hardware processor, whether the identified category includes a specified number of annotated documents;

based on a determination that the identified category does not include the specified number of annotated documents, generating, by the corpus generator and enricher that is executed by the at least one hardware processor, for the identified category, additional annotated documents, wherein the annotated documents and the additional annotated documents of the identified category together represent a corpus;

training, by an artificial intelligence model generator that is executed by the at least one hardware processor, using the corpus, an artificial intelligence based decision support model;

ascertaining, by an inquiry response generator that is executed by the at least one hardware processor, an inquiry related to an entity of the corpus; and

invoking, by the inquiry response generator that is executed by the at least one hardware processor, the artificial intelligence based decision support model to generate a response to the inquiry.

11. The method according to claim 10 , wherein generating, by the entity and relation annotator that is executed by the at least one hardware processor, based on the annotated training documents, the entity and relation annotation model further comprises:

transforming, by the entity and relation annotator that is executed by the at least one hardware processor, each annotated training document of the annotated training documents into a vector representation; and

generating, by the entity and relation annotator that is executed by the at least one hardware processor, based on vector representations of the annotated training documents, the entity and relation annotation model.

12. The method according to claim 10 , further comprising:

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, entities for which a difference between entity confidence scores is less than a specified numerical value;

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, relations for which a difference between relation confidence scores is less than the specified numerical value;

generating, by the entity and relation annotator that is executed by the at least one hardware processor, another inquiry for verification of the entities for which the difference between the entity confidence scores is less than the specified numerical value, and the relations for which the difference between the relation confidence scores is less than the specified numerical value; and

training, by the entity and relation annotator that is executed by the at least one hardware processor, based on a response to the other inquiry for verification of the entities for which the difference between the entity confidence scores is less than the specified numerical value, and the relations for which the difference between the relation confidence scores is less than the specified numerical value, the entity and relation annotation model.

13. The method according to claim 10 , further comprising:

categorizing, by the document categorizer that is executed by the at least one hardware processor, each document of the document set that is to be annotated into the respective category of the plurality of categories; and

identifying, by the entity and relation annotator that is executed by the at least one hardware processor, based on application of the entity and relation annotation model to documents of the identified category, entities and relations between the entities for each document of the identified category to generate annotated documents for the identified category.

14. The method according to claim 10 , wherein categorizing each annotated document of the annotated document set into the respective category of the plurality of categories further comprises:

transforming, by the document categorizer that is executed by the at least one hardware processor, each annotated document of the annotated document set into an entity vector;

grouping, by the document categorizer that is executed by the at least one hardware processor, based on the entity vector for each annotated document, semantically similar entities; and

categorizing, by the document categorizer that is executed by the at least one hardware processor, based on the grouping, each annotated document of the annotated document set into the respective category of the plurality of categories.

15. A non-transitory computer readable medium having stored thereon machine readable instructions, the machine readable instructions, when executed by at least one hardware processor, cause the at least one hardware processor to:

generate, based on annotated training documents, an entity and relation annotation model;

ascertain a document set that is to be annotated;

identify, based on application of the entity and relation annotation model to the document set, entities and relations between the entities for each document of the document set to generate an annotated document set;

determine, for each identified entity of the identified entities, an entity confidence score;

determine, for each identified relation of the identified relations between the entities, a relation confidence score;

identify, based on the entity confidence score, an entity that includes an entity confidence score that is less than an entity confidence score threshold;

identify, based on the relation confidence score, a relation that includes a relation confidence score that is less than a relation confidence score threshold;

generate another inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold;

train, based on a response to the other inquiry for verification of the entity and the relation that respectively include the entity confidence score and the relation confidence score that are respectively less than the entity confidence score threshold and the relation confidence score threshold, the entity and relation annotation model;

categorize each annotated document of the annotated document set into a respective category of a plurality of categories;

identify a category of the plurality of categories;

determine whether the identified category includes a specified number of annotated documents;

based on a determination that the identified category does not include the specified number of annotated documents, generate, for the identified category, additional annotated documents, wherein the annotated documents and the additional annotated documents of the identified category together represent a corpus;

train, using the corpus, an artificial intelligence based decision support model;

ascertain an inquiry related to an entity of the corpus; and

generate, by invoking the artificial intelligence based decision support model, a response to the inquiry.

16. The non-transitory computer readable medium according to claim 15 , wherein the machine readable instructions to generate, for the identified category, additional annotated documents, when executed by the at least one hardware processor, further cause the at least one hardware processor to:

segment the corpus into a plurality of sections to generate a preprocessed and segmented entity-annotated textual corpus.

17. The non-transitory computer readable medium according to claim 16 , wherein the machine readable instructions to generate, for the identified category, additional annotated documents, when executed by the at least one hardware processor, further cause the at least one hardware processor to:

transform each segment of the preprocessed and segmented entity-annotated textual corpus into a character representation to generate a plurality of character representations;

consolidate, based on the character representations, segments of the preprocessed and segmented entity-annotated textual corpus; and

learn, from each segment of the segments, character embeddings.

18. The non-transitory computer readable medium according to claim 17 , wherein the machine readable instructions to generate, for the identified category, additional annotated documents, when executed by the at least one hardware processor, further cause the at least one hardware processor to:

ascertain a seed document;

segment the seed document into another plurality of sections;

transform each segment of the segmented seed document into character represented vector embeddings;

generate a plurality of corresponding segments specific to each transformed segment of the segmented seed document; and

generate, based on the plurality of corresponding segments and the learned character embeddings, the additional annotated documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2018
From: GUGGILLA, CHINNAPPA; SHISHTLA, PRANEETH; SHIVARAM, MADHURA
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 047756/0135 →
Priority Claims (1)
IN 201811032723 · Aug 31, 2018 · national
Continuity (1)
Related Publication 20200073882A1 · Mar 5, 2020