IP Library Granted Patent US 11,037,549
Granted Patent B1
US 11,037,549 · App. 16/929,675 · Granted Jun 15, 2021

System and method for automating the training of enterprise customer response systems using a range of dynamic or generic data sets

Inventors: Santosh Kulkarni (Glen Waverley, AU); Sachin Pathiyan Cherumanal (Glen Huntly, AU); Damiano Spina (Elsternwick, AU); Minyi Li (Burwood, AU); Lawrence Cavedon (Fitzroy, AU)
Assignee: Inference Communications Pty Ltd
G10L15/063G06N20/00G10L15/10G10L15/183G10L15/1815G10L15/22G10L15/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,037,549
App. No.
16/929,675
Granted
Jun 15, 2021
Kind
B1
Abstract

A system and method for automating the training of enterprise customer response systems using a range of dynamic or generic data sets, used to gradually take human supervision and intervention out of the training process for enterprise WA and similar automated response engines, by training existing machine learning models or engines using a heuristic middleman annotation assistant that helps map generic/public/new datasets to the existing machine learning model or engine, and allowing for limited human oversight over the remaining unknown or badly classified data segments which is used to further teach the heuristic and classification model until the human oversight is no longer needed for the heuristic to learn and map newer datasets, thereby reducing human, dollar, and time costs, and improving automated response system efficiency.

Claims (36)

1. A system for automating the training of enterprise virtual agents using a range of dynamic or generic data sets, comprising:

an annotation assistant comprising at least a machine learning model, a first plurality of programming instructions stored in a memory of, and operating on at least one processor of, a computing device, wherein the first plurality of programming instructions, when operating on the at least one processor, cause the computing device to:

receive data from a plurality of verbal utterance datasets, wherein the data comprises a plurality of verbal utterances;

embed a plurality of user intention vectors in the received data to produce an utterance model, wherein each of the plurality of embedded user intention vectors is associated with a verbal utterance and is produced using previously-stored association data between an utterance and an intention;

receive annotated new verbal utterance data from an uncertainty classifier;

calculate a confidence index for the annotated new verbal utterance data, wherein the confidence index indicates a degree of confidence in the annotations in the annotated new verbal utterance data;

if the confidence index is above a threshold value, incorporate the annotated new verbal utterance data into the utterance model; and

send the utterance model to an automated virtual agent; and

an uncertainty classifier comprising at least a second plurality of programming instructions stored in a memory of, and operating on at least one processor of, a computing device, wherein the second plurality of programming instructions, when operating on the at least one processor, cause the computing device to:

receive new verbal utterance data from an automated virtual agent;

compare known verbal utterance data and user intention vectors, to new verbal utterance data with unknown or unclassified user intention vectors, the comparison comprising a plurality of semantic comparison operations;

use a heuristic to derive similarity and uncertainty index values for the unknown or unclassified user intention vectors;

annotate at least a portion of the new verbal utterance data, to embed a plurality of user intention vectors, wherein for each of the plurality of user intention vectors the similarity index value is greater than the uncertainty index value; and

send the annotated new verbal utterance data to the annotation assistant.

2. The system of claim 1 , wherein a semantic filter operates as part of an annotation assistant, and receives verbal utterance data from an automated virtual agent that failed to be classified with a user intention vector, and assigns a user intention vector to the verbal utterance data based on the semantic similarities to other, properly assigned verbal utterances.

3. The system of claim 1 , wherein the annotation assistant trains and creates a new instance of the automated virtual agent after new verbal utterance data is embedded with user intention vectors, thereby training the new instance of the automated virtual agent with the newly embedded data.

4. The system of claim 1 , wherein the verbal utterance datasets are generalized and not designed or curated specifically for any automated virtual agent or annotation system.

5. The system of claim 1 , wherein the plurality of semantic comparison operations includes the use of information distance modeling to cluster at least a portion of the new verbal utterance data with a plurality of the unknown or unclassified user intention vectors.

6. The system of claim 5 , wherein the information distance modeling comprises the use of Euclidian distance.

7. A method for automating the training of enterprise customer response systems using a range of dynamic or generic data sets, comprising the steps of:

receiving data from a plurality of verbal utterance datasets, using an annotation assistant;

embedding a plurality of user intention vectors in the received data to produce an utterance model, wherein each of the plurality of embedded user intention vectors is associated with a verbal utterance and is produced using previously-stored association data between an utterance and an intention;

receiving annotated new verbal utterance data from an uncertainty classifier, using the annotation assistant;

calculating a confidence index for the annotated new verbal utterance data, wherein the confidence index indicates a degree of confidence in the annotations in the annotated new verbal utterance data, using the annotation assistant;

if the confidence index is above a threshold value, incorporating the annotated new verbal utterance data into the utterance model, using the annotation assistant;

sending the utterance model to an automated virtual agent, using the annotation assistant;

receiving new verbal utterance data from an automated virtual agent, using an uncertainty classifier;

comparing known verbal utterance data and user intention vectors, to new verbal utterance data with unknown or unclassified user intention vectors, the comparison comprising a plurality of semantic comparison operations, using an uncertainty classifier;

using a heuristic to derive similarity and uncertainty index values for the unknown or unclassified user intention vectors, using the uncertainty classifier;

annotating at least a portion of the new verbal utterance data, to embed a plurality of user intention vectors, wherein for each of the plurality of user intention vectors the similarity index value is greater than the uncertainty index value;

sending the annotated new verbal utterance data to the annotation assistant.

8. The method of claim 7 , wherein a semantic filter operates as part of an annotation assistant, and receives verbal utterance data from an automated virtual agent that failed to be classified with a user intention vector, and assigns a user intention vector to the verbal utterance data based on the semantic similarities to other, properly assigned verbal utterances.

9. The method of claim 7 , wherein the annotation assistant trains and creates a new instance of the automated virtual agent after new verbal utterance data is embedded with user intention vectors, thereby training the new instance of the automated virtual agent with the newly embedded data.

10. The method of claim 7 , wherein the verbal utterance datasets are generalized and not designed or curated specifically for any automated virtual agent or annotation system.

11. The method of claim 7 , wherein the plurality of semantic comparison operations includes the use of information distance modeling to cluster at least a portion of the new verbal utterance data with a plurality of the unknown or unclassified user intention vectors.

12. The method of claim 11 , wherein the information distance modeling comprises the use of Euclidian distance.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2020
From: INFERENCE SOLUTIONS PTY LTD
To: INFERENCE COMMUNICATIONS PTY LTD
Reel/Frame 054033/0597 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2020
From: KULKARNI, SANTOUSH
To: INFERENCE SOLUTIONS PTY LTD
Reel/Frame 054022/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2020
From: ROYAL MELBOURNE INSTITUTE OF TECHNOLOGY
To: INFERENCE SOLUTIONS PTY LTD
Reel/Frame 054022/0259 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2020
From: CHERUMANAL, SACHIN PATHIYAN; SPINA, DAMIANO; LI, MINYI; CAVEDON, LAWRENCE
To: ROYAL MELBOURNE INSTITUTE OF TECHNOLOGY
Reel/Frame 053968/0817 →
Continuity (1)
Provisional Application 63035506 · Jun 5, 2020
Cited By (7)
US 12,197,876 US 12,314,675 US 12,340,358 US 12,437,158 US 12,619,830 US 12,657,832 US 12,688,440