IP Library Granted Patent US 12,118,986
Granted Patent B2
US 12,118,986 · App. 17/380,425 · Granted Oct 15, 2024

System and method for automated processing of natural language using deep learning model encoding

Inventor: Rijul Magu (Morrisville, NC)
Assignee: Conduent Business Services, LLC
G10L15/1815G06N3/08G10L15/04G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,118,986
App. No.
17/380,425
Granted
Oct 15, 2024
Kind
B2
Abstract

Automated systems and methods are provided for processing natural language, comprising obtaining first and second digitally-encoded speech representations, respectively corresponding to an agent script for and a voice recording of a telecommunication interaction; generating a similarity structure based on the speech representations, the similarity structure representing a degree of semantic similarity between the speech representations; matching markers in the first speech representation to markers in the second speech representation based on the similarity structure; and dividing the telecommunication interaction into a plurality of sections based on the matching.

Claims (51)

1. A computer-implemented method for processing natural language, comprising:

obtaining a first digitally-encoded speech representation corresponding to an agent script for a telecommunication interaction;

obtaining a second digitally-encoded speech representation corresponding to a voice recording of the telecommunication interaction;

generating a similarity structure based on the first digitally-encoded speech representation and the second digitally-encoded speech representation, the similarity structure representing a degree of semantic similarity between the first digitally-encoded speech representation and the second digitally-encoded speech representation;

matching at least one first marker in the first digitally-encoded speech representation to at least one second marker in the second digitally-encoded speech representation based on the similarity structure, wherein matching further comprises identifying discrepancies between an order of markers of first digitally-encoded speech representation and an order of markers of second digitally-encoded speech representation; and

dividing the telecommunication interaction into a plurality of sections based on the matching.

2. The method according to claim 1 , wherein obtaining the first digitally-encoded speech representation includes applying a machine learning algorithm to the agent script to generate the first digitally-encoded speech representation, wherein the first digitally-encoded speech representation is a first numeric vector having a predetermined size.

3. The method according to claim 2 , wherein obtaining the second digitally-encoded speech representation includes:

generating a transcript from the voice recording; and

applying the machine learning algorithm to the transcript to generate the second digitally-encoded speech representation, wherein the first digitally-encoded speech representation is a second numeric vector having the predetermined size.

4. The method according to claim 3 , wherein generating the similarity structure includes applying a vector operation to the first numeric vector and the second numeric vector, thereby to generate a similarity matrix including a plurality of matrix elements.

5. The method according to claim 4 , wherein matching the at least one first marker to the at least one second marker includes:

comparing the plurality of matrix elements to a predetermined similarity threshold value; and

discarding respective ones of the plurality of matrix elements which are less than the similarity threshold value.

6. The method according to claim 5 , wherein dividing the telecommunication interaction into the plurality of sections includes:

applying a clustering algorithm to each of the plurality of matrix elements which have not been discarded; and

generating at least one section start timestamp and at least one section stop timestamp.

7. The method according to claim 1 , wherein the first digitally-encoded speech representation includes a first plurality of embeddings respectively corresponding to a plurality of expected utterances in the telecommunication interaction, and the second digitally-encoded speech representation includes a second plurality of embeddings respectively corresponding to a plurality of actual utterances in the telecommunication interaction.

8. The method according to claim 7 , wherein identifying discrepancies between the order of markers of the first digitally-encoded speech representation and the order of markers of the second digitally-encoded speech representation further comprises:

identifying which of the plurality of expected utterances do not have a corresponding one of the plurality of actual utterances; and

identifying which of the plurality of expected utterances have multiple corresponding ones of the plurality of actual utterances.

9. The method according to claim 7 , wherein

the order of markers of first digitally-encoded speech representation comprises an expected order of the plurality of expected utterances,

the order of markers of second digitally-encoded speech representation comprises an actual order of the plurality of actual utterances, and

the method further comprises identifying discrepancies between the expected order and the actual order.

10. The method according to claim 1 , further comprising calculating a duration of at least one of the plurality of sections.

11. A computing system for processing natural language, the system comprising:

at least one electronic processor; and

a non-transitory computer-readable medium storing instructions that, when executed by the at least one electronic processor, cause the at least one electronic processor to perform operations comprising:

obtaining a first digitally-encoded speech representation corresponding to an agent script for a telecommunication interaction,

obtaining a second digitally-encoded speech representation corresponding to a voice recording of the telecommunication interaction,

generating a similarity structure based on the first digitally-encoded speech representation and the second digitally-encoded speech representation, the similarity structure representing a degree of semantic similarity between the first digitally-encoded speech representation and the second digitally-encoded speech representation,

matching at least one first marker in the first digitally-encoded speech representation to at least one second marker in the second digitally-encoded speech representation based on the similarity structure, wherein matching further comprises identifying discrepancies between an order of markers of first digitally-encoded speech representation and an order of markers of second digitally-encoded speech representation, and

dividing the telecommunication interaction into a plurality of sections based on the matching.

12. The system according to claim 11 , wherein obtaining the first digitally-encoded speech representation includes applying a machine learning algorithm to the agent script to generate the first digitally-encoded speech representation, wherein the first digitally-encoded speech representation is a first numeric vector having a predetermined size.

13. The system according to claim 12 , wherein obtaining the second digitally-encoded speech representation includes:

generating a transcript from the voice recording; and

applying the machine learning algorithm to the transcript to generate the second digitally-encoded speech representation, wherein the first digitally-encoded speech representation is a second numeric vector having the predetermined size.

14. The system according to claim 13 , wherein generating the similarity structure includes applying a vector operation to the first numeric vector and the second numeric vector, thereby to generate a similarity matrix including a plurality of matrix elements.

15. The system according to claim 14 , wherein matching the at least one first marker to the at least one second marker includes:

comparing the plurality of matrix elements to a predetermined similarity threshold value; and

discarding respective ones of the plurality matrix elements which are less than the similarity threshold value.

16. The system according to claim 15 , wherein dividing the telecommunication interaction into the plurality of sections includes:

applying a clustering algorithm to each of the plurality of matrix elements which have not been discarded; and

generating at least one section start timestamp and at least one section stop timestamp.

17. The system according to claim 11 , wherein obtaining the first digitally-encoded speech representation includes retrieving a predetermined numeric vector from a memory of the system, the numeric vector including a plurality of elements respectively corresponding to subsections of the agent script.

18. The system according to claim 11 , wherein the first digitally- encoded speech representation includes a first plurality of embeddings respectively corresponding to a plurality of expected utterances in the telecommunication interaction, and the second digitally-encoded speech representation includes a second plurality of embeddings respectively corresponding to a plurality of actual utterances in the telecommunication interaction.

19. The system according to claim 18 , wherein identifying discrepancies between the order of markers of the first digitally-encoded speech representation and the order of markers of the second digitally-encoded speech representation further comprises:

identifying which of the plurality of expected utterances do not have a corresponding one of the plurality of actual utterances; and

identifying which of the plurality of expected utterances have multiple corresponding ones of the plurality of actual utterances.

20. The system according to claim 11 , the operations further comprising calculating a duration of at least one of the plurality of sections.

Assignments (3)
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: U.S. BANK, NATIONAL ASSOCIATION
Reel/Frame 057969/0445 →
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 057970/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2021
From: MAGU, RIJUL
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 056927/0289 →
Continuity (1)
Related Publication 20230025114A1 · Jan 26, 2023