IP Library › Granted Patent US 12,106,056
Granted Patent B2
US 12,106,056 · App. 17/745,230 · Granted Oct 1, 2024

Natural language processing of structured interactions

Inventors: Jason T. Lam (Scarborough, CA); Joel D. Young (Milpitas, CA); Honglei Zhuang (Santa Clara, CA); Netanel Weizman (San Jose, CA); Scott C. Parish (Santa Monica, CA); Bryan N. Cavanagh (Wheat Ridge, CO); Christopher C. Cole (Santa Monica, CA); Alycia M. Damp (Toronto, CA); Eszter Fodor (London, GB); Laura A. Kruizenga (Hennepin, MN); Janice C. Oh (San Leandro, CA); Rachel M. Policastro (Denver, CO); Geoffrey C. Thomas (Denver, CO); Gregory A. Walloch (Grass Valley, CA)
Assignee: HIREGUIDE PBC
G06F40/35G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,106,056
App. No.
17/745,230
Granted
Oct 1, 2024
Kind
B2
Abstract

One embodiment of the present invention sets forth a technique for analyzing a transcript of a structured interaction. The technique includes determining a first portion of the transcript that corresponds to a first dialogue act. The technique also includes matching the first portion of the transcript to a first component of a script for the structured interaction based on a first set of embeddings for the first portion of the transcript and a second set of embeddings associated with the first component of the script. The technique further includes causing a first mapping between the first portion of the transcript and the first component to be outputted.

Claims (61)

1. A computer-implemented method for analyzing a transcript of a structured interaction, the method comprising:

determining that a first portion of the transcript corresponds to a first dialogue act;

matching the first portion of the transcript to a first component of a script for the structured interaction by:

generating a first term frequency-inverse document frequency (TF-IDF) vector for the first portion of the transcript based on a vocabulary associated with the transcript;

executing an embedding model that generates a first embedding representing a sequence of words in the first portion of the transcript;

aggregating a plurality of embeddings for a plurality of cues in the script into a second embedding for an attribute that is associated with the plurality of cues and corresponds to the first component of the script;

aggregating a plurality of TF-IDF vectors for the plurality of cues into a second TF-IDF vector for the attribute;

aggregating (i) a first similarity between the first TF-IDF vector and the second TF-IDF vector and (ii) a second similarity between the first embedding and the second embedding for the first component of the script into an overall similarity between the first portion of the transcript and the first component of the script; and

matching the first portion of the transcript to the first component of the script based on the overall similarity;

outputting, in a user interface, a first mapping between the first portion of the transcript and the first component of the script;

receiving, via the user interface, user input associated with the first mapping; and

training the embedding model based on the user input.

2. The computer-implemented method of claim 1 , wherein determining the first portion of the transcript that corresponds to the first dialogue act comprises:

executing an encoder neural network that converts the first portion of the transcript into one or more embeddings; and

applying one or more classification layers to the one or more embeddings to determine the first dialogue act associated with the first portion of the transcript.

3. The computer-implemented method of claim 2 , wherein determining the first portion of the transcript that corresponds to the first dialogue act further comprises determining a sentence within the transcript that corresponds to the first portion.

4. The computer-implemented method of claim 2 , wherein the first dialogue act comprises an interrogatory.

5. The computer-implemented method of claim 1 , wherein the embedding model comprises at least one of a Bidirectional Encoder Representations from Transformers (BERT) model or a bidirectional gated recurrent unit (GRU).

6. The computer-implemented method of claim 1 , wherein the plurality of cues comprises at least one of a question, a prompt, or a topic included in the first component of the script.

7. The computer-implemented method of claim 1 , wherein

the attribute comprises a skill.

8. The computer-implemented method of claim 1 , further comprising:

determining a lack of match between a portion of the transcript and a second component of the script; and

outputting an indication that the second component of the script is not included in the structured interaction.

9. The computer-implemented method of claim 1 , wherein outputting the first mapping between the first portion of the transcript and the first component of the script comprises outputting the first portion of the transcript, the first component of the script, and a recording of the first portion of the transcript within the user interface.

10. The computer-implemented method of claim 1 , wherein the overall similarity is computed using at least one of a sum, a weighted sum, or an average of the first similarity and the second similarity.

11. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

determining that a first portion of a transcript corresponds to a first dialogue act;

matching the first portion of the transcript to a first component of a script for a structured interaction by:

generating a first term frequency-inverse document frequency (TF-IDF) vector for the first portion of the transcript based on a vocabulary associated with the transcript;

executing an embedding model that generates a first embedding representing a sequence of words in the first portion of the transcript;

aggregating a plurality of embeddings for a plurality of cues in the script into a second embedding for an attribute that is associated with the plurality of cues and corresponds to the first component of the script;

aggregating a plurality of TF-IDF vectors for the plurality of cues into a second TF-IDF vector for the attribute;

aggregating (i) a first similarity between the first TF-IDF vector and the second TF-IDF vector and (ii) a second similarity between the first embedding and the second embedding for the first component of the script into an overall similarity between the first portion of the transcript and the first component of the script; and

matching the first portion of the transcript to the first component of the script based on the overall similarity;

outputting, in a user interface, a first mapping between the first portion of the transcript and the first component of the script;

receiving, via the user interface, user input associated with the first mapping; and

training the embedding model based on the user input.

12. The one or more non-transitory computer readable media of claim 11 , wherein the operations further comprise:

outputting a first indication of the first dialogue act in association with the first portion of the transcript; and

outputting a second indication of a second dialogue act in association with a second portion of the transcript that follows the first dialogue act in the transcript.

13. The one or more non-transitory computer readable media of claim 12 , wherein the first dialogue act comprises an interrogatory and the second dialogue act comprises an answer.

14. The one or more non-transitory computer readable media of claim 11 , wherein

the vocabulary comprises a set of words included in the transcript.

15. The one or more non-transitory computer readable media of claim 11 , wherein

the first similarity and

the second similarity comprise at least one of a dot product, a cosine similarity, or a Euclidean distance.

16. A system, comprising:

one or more memories that store instructions, and

one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to:

determine that a first portion of a transcript corresponds to an interrogatory dialogue act;

match the first portion of the transcript to a first component of a script for a structured interaction by:

generating a first term frequency-inverse document frequency (TF-IDF) vector for the first portion of the transcript based on a vocabulary associated with the transcript;

executing an embedding model that generates a first embedding representing a sequence of words in the first portion of the transcript;

aggregating a plurality of embeddings for a plurality of cues in the script into a second embedding for an attribute that is associated with the plurality of cues and corresponds to the first component of the script;

aggregating a plurality of TF-IDF vectors for the plurality of cues into a second TF-IDF vector for the attribute;

aggregating (i) a first similarity between the first TF-IDF vector and the second TF-IDF vector and (ii) a second similarity between the first embedding and the second embedding for the first component of the script into an overall similarity between the first portion of the transcript and the first component of the script; and

matching the first portion of the transcript to the first component of the script based on the overall similarity;

output, in a user interface, a first mapping between the first portion of the transcript and the first component of the script;

receive, via the user interface, user input associated with the first mapping; and

train the embedding model based on the user input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: LAM, JASON T.; YOUNG, JOEL D.; ZHUANG, HONGLEI; WEIZMAN, NETANEL; PARISH, SCOTT C.; CAVANAGH, BRYAN N.; COLE, CHRISTOPHER C.; DAMP, ALYCIA M.; FODOR, ESZTER; KRUIZENGA, LAURA A.; OH, JANICE C.; POLICASTRO, RACHEL M.; THOMAS, GEOFFREY C.; WALLOCH, GREGORY A.
To: HIREGUIDE PBC
Reel/Frame 059920/0837 →
Continuity (2)
Provisional Application 63285976 · Dec 3, 2021
Related Publication 20230177275A1 · Jun 8, 2023
Cited By (2)
US 12,368,938 US 12,380,285