IP Library Granted Patent US 11,410,644
Granted Patent B2
US 11,410,644 · App. 16/996,653 · Granted Aug 9, 2022

Generating training datasets for a supervised learning topic model from outputs of a discovery topic model

Inventors: Michael McCourt (Santa Barbara, CA); Anoop Praturu (Santa Barbara, CA)
Assignee: INVOCA, INC.
G10L15/1815G06N5/04G06N20/00G10L15/063G10L15/197G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,644
App. No.
16/996,653
Granted
Aug 9, 2022
Kind
B2
Abstract

Systems and methods for generating training data for a supervised topic modeling system from outputs of a topic discovery model are described herein. In an embodiment, a system receives a plurality of digitally stored call transcripts and, using a topic model, generates an output which identifies a plurality of topics represented in the plurality of digitally stored call transcripts. Using the output of the topic model, the system generates an input dataset for a supervised learning model by identify a first subset of the plurality of digitally stored call transcripts that include the particular topic, storing a positive value for the first subset, identifying a second subset that do not include the particular topic, and storing a negative value for the second subset. The input training dataset is then used to train a supervised learning model.

Claims (54)

1. A computer system comprising:

one or more processors;

a memory coupled to the one or more processors and storing sequences of instructions which, when executed by the one or more processors, causes the one or more processors to perform:

receiving a plurality of digitally stored call transcripts that have been prepared from digitally recorded voice calls;

using a topic model of an artificial intelligence machine learning system, the topic model being programmed to model words of a call as a function of one or more word distributions for each topic of a plurality of topics, generating an output of the topic model which identifies the plurality of topics represented in the plurality of digitally stored call transcripts;

using the output of the topic model, generating an input dataset for a supervised learning model by performing:

identifying, based, at least in part, on the output of the topic model, a first subset of the plurality of digitally stored call transcripts that include a particular topic;

identifying, based, at least in part, on the output of the topic model, a second subset of the plurality of digitally stored call transcripts that do not include the particular topic;

storing a positive output value for the first subset of the plurality of digitally stored call transcripts;

storing a negative output value for the second subset of the plurality of digitally stored call transcript;

determining that a second topic of the plurality of topics corresponds to the particular topic by: determining that a first threshold number of a second threshold number of highest probability words in a particular probability distribution for the particular topic are in the second threshold number of highest probability words in a second probability distribution for the second topic; identifying one or more digitally stored call transcripts of the second subset of the plurality of digitally stored call transcripts that include the second topic; removing the one or more digitally stored call transcripts from the input dataset; and

training the supervised learning model using the generated input dataset, the supervised learning model being configured to compute, for a new digitally stored call transcript, a probability that the particular topic was discussed during a digitally recorded voice call corresponding to the new digitally stored call transcript.

2. The computer system of claim 1 , wherein identifying the first subset of the plurality of digitally stored call transcripts comprises determining, for each digitally stored call transcript of the first subset of the plurality of digitally stored call transcripts, that a proportion of words in the digitally stored call transcript that exist in the particular topic is greater than or equal to a threshold value.

3. The computer system of claim 2 , wherein the threshold value comprises a first value corresponding to an a priori expected proportion of words in the digitally stored call transcript that are drawn from the particular topic multiplied by a second value corresponding to a length of a digitally recorded voice call corresponding to the digitally stored call transcript.

4. The computer system of claim 2 , wherein identifying the first subset of the plurality of digitally stored call transcripts further comprises:

drawing a sample from a probability distribution of words for the particular topic;

determining, for each digitally stored call transcript of the first subset of the plurality of digitally stored call transcripts, that a number of words in the digitally stored call transcript that correspond to the sample is greater than a second threshold value.

5. The computer system of claim 1 , wherein identifying the second subset of the plurality of digitally stored call transcripts comprises determining, for each digitally stored call transcript of the second subset of the plurality of digitally stored call transcripts, that a proportion of words in the digitally stored call transcript that exist in the particular topic is less than a threshold value.

6. The computer system of claim 5 , wherein the threshold value comprises a first value corresponding to an a priori expected proportion of words in the digitally stored call transcript that are drawn from the particular topic multiplied by a second value corresponding to a length of a digitally recorded voice call corresponding to the digitally stored call transcript.

7. The computer system of claim 5 , wherein identifying the second subset of the plurality of digitally stored call transcripts further comprises:

drawing a sample from a probability distribution of words for the particular topic;

determining, for each digitally stored call transcript of the second subset of the plurality of digitally stored call transcripts, that a number of words in the digitally stored call transcript that correspond to the sample equals zero.

8. The computer system of claim 1 , wherein generating the input dataset for the supervised learning model further comprises performing:

identifying, based, at least in part, on the output of the topic model, a third subset of the plurality of digitally stored call transcripts that do not meet stored criteria for the first subset of the plurality of digitally stored call transcripts and the second subset of the plurality of digitally stored call transcripts;

removing the third subset of the plurality of digitally stored call transcripts from the input dataset.

9. A method comprising:

receiving a plurality of digitally stored call transcripts that have been prepared from digitally recorded voice calls;

using a topic model of an artificial intelligence machine learning system, the topic model modeling words of a call as a function of one or more word distributions for each topic of a plurality of topics, generating an output of the topic model which identifies the plurality of topics represented in the plurality of digitally stored call transcripts;

using the output of the topic model, generating an input dataset for a supervised learning model by performing:

identifying, based, at least in part, on the output of the topic model, a first subset of the plurality of digitally stored call transcripts that include a particular topic;

identifying, based, at least in part, on the output of the topic model, a second subset of the plurality of digitally stored call transcripts that do not include the particular topic;

storing a positive output value for the first subset of the plurality of digitally stored call transcripts;

storing a negative output value for the second subset of the plurality of digitally stored call transcript;

determining that a second topic of the plurality of topics corresponds to the particular topic by: determining that a first threshold number of a second threshold number of highest probability words in a particular probability distribution for the particular topic are in the second threshold number of highest probability words in a second probability distribution for the second topic; identifying one or more digitally stored call transcripts of the second subset of the plurality of digitally stored call transcripts that include the second topic; removing the one or more digitally stored call transcripts from the input dataset; and

training the supervised learning model using the generated input dataset, wherein the supervised learning model is configured to compute, for a new digitally stored call transcript, a probability that the particular topic was discussed during a digitally recorded voice call corresponding to the new digitally stored call transcript.

10. The method of claim 1 , wherein determining that the second topic corresponds to the particular topic comprises:

determining that a subset of the plurality of topics comprise scripted topics;

only determining that the second topic corresponds to the particular topic if the second topic is not in the subset of the plurality of topics.

11. The method of claim 9 , wherein determining that the second topic corresponds to the particular topic comprises:

determining that a subset of the plurality of topics comprise scripted topics;

only determining that the second topic corresponds to the particular topic if the second topic is not in the subset of the plurality of topics.

12. The method of claim 9 , wherein identifying the first subset of the plurality of digitally stored call transcripts comprises determining, for each digitally stored call transcript of the first subset of the plurality of digitally stored call transcripts, that a proportion of words in the digitally stored call transcript that exist in the particular topic is greater than or equal to a threshold value.

13. The method of claim 12 , wherein the threshold value comprises a first value corresponding to an a priori expected proportion of words in the digitally stored call transcript that are drawn from the particular topic multiplied by a second value corresponding to a length of a digitally recorded voice call corresponding to the digitally stored call transcript.

14. The method of claim 12 , wherein identifying the first subset of the plurality of digitally stored call transcripts further comprises:

drawing a sample from a probability distribution of words for the particular topic;

determining, for each digitally stored call transcript of the first subset of the plurality of digitally stored call transcripts, that a number of words in the digitally stored call transcript that correspond to the sample is greater than a second threshold value.

15. The method of claim 11 , wherein identifying the second subset of the plurality of digitally stored call transcripts comprises determining, for each digitally stored call transcript of the second subset of the plurality of digitally stored call transcripts, that a proportion of words in the digitally stored call transcript that exist in the particular topic is less than a threshold value.

16. The method of claim 15 , wherein the threshold value comprises a first value corresponding to an a priori expected proportion of words in the digitally stored call transcript that are drawn from the particular topic multiplied by a second value corresponding to a length of a digitally recorded voice call corresponding to the digitally stored call transcript.

17. The method of claim 15 , wherein identifying the second subset of the plurality of digitally stored call transcripts further comprises:

drawing a sample from a probability distribution of words for the particular topic;

determining, for each digitally stored call transcript of the second subset of the plurality of digitally stored call transcripts, that a number of words in the digitally stored call transcript that correspond to the sample equals zero.

18. The method of claim 11 , wherein generating the input dataset for the supervised learning model further comprises performing:

identifying, based, at least in part, on the output of the topic model, a third subset of the plurality of digitally stored call transcripts that do not meet stored criteria for the first subset of the plurality of digitally stored call transcripts and the second subset of the plurality of digitally stored call transcripts;

removing the third subset of the plurality of digitally stored call transcripts from the input dataset.

Assignments (5)
SECURITY INTEREST Recorded Aug 6, 2024
From: INVOCA, INC.
To: BANC OF CALIFORNIA (FORMERLY KNOWN AS PACIFIC WESTERN BANK)
Reel/Frame 068200/0412 →
RELEASE OF SECURITY INTEREST Recorded Jan 24, 2023
From: ORIX GROWTH CAPITAL, LLC
To: INVOCA, INC.
Reel/Frame 062463/0390 →
REAFFIRMATION OF AND SUPPLEMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jan 28, 2022
From: INVOCA, INC.
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 058892/0404 →
REAFFIRMATION OF AND SUPPLEMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 21, 2021
From: INVOCA, INC.
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 057884/0947 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2020
From: MCCOURT, MICHAEL; PRATURU, ANOOP
To: INVOCA, INC.
Reel/Frame 053551/0087 →
Continuity (3)
Provisional Application 62980092 · Feb 21, 2020
Provisional Application 62923323 · Oct 18, 2019
Related Publication 20210118432A1 · Apr 22, 2021
Cited By (2)
US 12,322,382 US 12,499,881