IP Library Granted Patent US 11,227,584
Granted Patent B2
US 11,227,584 · App. 16/780,340 · Granted Jan 18, 2022

System and method for determining the compliance of agent scripts

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,584
App. No.
16/780,340
Filed
Feb 3, 2020
Granted
Jan 18, 2022
Kind
B2
Art Unit
2663
USPC
704/239
Abstract

Systems and methods of script identification in audio data obtained from audio data. The audio data is segmented into a plurality of utterances. A script model representative of a script text is obtained. The plurality of utterances are decoded with the script model. A determination is made if the script text occurred in the audio data.

Claims (70)

1. A method of script identification in audio data, the method comprising:

receiving audio data, wherein the audio data includes speech from a specific customer service agent;

performing voice activity detection to segment the audio data into a plurality of utterances, wherein each utterance is a segment of speech separated by a segment of non-speech;

performing speech analytics on each of the plurality of utterances, wherein the speech analytics includes:

filtering the plurality of utterances into a subset of utterances attributed to the specific customer service agent, and

identifying at least one acoustic feature for each of the utterances in the subset of utterances, wherein the at least one acoustic feature is used to distinguish phonemes in each of the utterances in the subset of utterances;

decoding each of the subset of the plurality of utterances to determine whether any of a plurality of script texts are contained in the utterance, wherein the decoding includes:

receiving at least one script model compilation, wherein each script model compilation includes text from a script text, a plurality of acceptable variations for the text of the script text, and at least one speaker model, wherein the at least one speaker model is specific to the specific customer service agent, further wherein the at least one speaker model is an acoustic model for the specific customer service agent,

comparing each script model compilation to the subset of the plurality of utterances using the identified acoustic features, and

determining the script model compilations present in each utterance of the subset of the plurality of utterances;

identifying each script text corresponding to each determined script model compilation;

determining compliance of the plurality of utterances, wherein determining compliance includes:

transcribing each utterance containing the script to produce an utterance transcript,

comparing the script text associated with each utterance transcript to determine a word error rate in the utterance transcript, and

determining the compliance of the utterance transcript based on the word error rate and a error threshold for the script text; and

initiating at least one remedial action if the utterance transcript is non-compliant to present on screen guidance to the specific customer service agent on a graphical display.

2. The method of claim 1 , wherein the audio data is real-time streamed audio data.

3. The method of claim 1 , wherein the remedial action is presented to the specific customer service agent in real-time.

4. The method of claim 1 , wherein the audio data is an interaction between at least the specific customer service agent and at least one customer.

5. The method of claim 1 , wherein at least one of the plurality of utterances consists of more than a single word.

6. A method of script identification in audio data, the method comprising:

receiving audio data;

performing voice activity detection to segment the audio data into a plurality of utterances, wherein each utterance is a segment of speech separated by a segment of non-speech;

performing speech analytics on each of the plurality of utterances, wherein the speech analytics includes:

filtering the plurality of utterances into a subset of utterances attributed to any customer service agents, and

identifying at least one acoustic feature for each of the utterances in the subset of utterances, wherein the at least one acoustic feature is used to distinguish phonemes in each of the utterances in the subset of utterances;

decoding each of the subset of the plurality of utterances to determine whether any of a plurality of script texts are contained in the utterance, wherein the decoding includes:

receiving at least one script model compilation, wherein each script model compilation includes text from a script text and a plurality of acceptable variations for the text of the script text,

comparing each script model compilation to the subset of the plurality of utterances using the identified acoustic features, and

determining the script model compilations present in each utterance of the subset of the plurality of utterances;

identifying each script text corresponding to each determined script model compilation;

determining compliance of the plurality of utterances with the corresponding identified script text, wherein determining compliance includes:

transcribing each utterance containing the script to produce an utterance transcript,

comparing the script text associated with each utterance transcript to determine a word error rate in the utterance transcript, and

determining the compliance of the utterance transcript based on the word error rate and an error threshold for the script text; and

initiating at least one remedial action if the utterance transcript is non-compliant to present on screen guidance to a customer service agent on a graphical display.

7. The method of claim 6 , wherein the audio data is real-time streamed audio data.

8. The method of claim 6 , wherein the audio data is an interaction between at least a customer service agent and at least one customer.

9. The method of claim 6 , wherein at least one of the plurality of utterances consists of more than a single word.

10. A method of script identification in audio data, the method comprising:

receiving audio data, wherein the audio data includes speech from a specific customer service agent;

performing voice activity detection to segment the audio data into a plurality of utterances, wherein each utterance is a segment of speech separated by a segment of non-speech;

filtering the plurality of utterances into a subset of utterances attributed to the specific customer service agent;

performing speech analytics on the subset of the plurality of utterances to identify at least one event, wherein the event is a detection of at least one keyword in the subset of the plurality of utterances, further wherein the identified event is associated with a requirement that the specific customer service agent speak at least one of a plurality of script texts in the audio data;

determining the least one script text required to be spoken by the specific customer service agent based on the identified at least one event;

receiving at least one script model compilation associated with the at least one script text, wherein each script model compilation includes text from the associated script text, a plurality of acceptable variations for the text of the associated script text, and at least one speaker model, wherein the at least one speaker model is specific to the specific customer service agent, further wherein the at least one speaker model is an acoustic model for the specific customer service agent;

determining compliance of the plurality of utterances to each script model compilation, wherein determining compliance includes:

transcribing the subset of the plurality of utterances to produce an utterance transcript,

comparing the script model compilation to the utterance transcript to determine a word error rate in the utterance transcript, and

determining the compliance of the utterance transcript based on the word error rate and an error threshold for the script model compilation; and

initiating at least one remedial action if the utterance transcript is non-compliant to present on screen guidance to the specific customer service agent on a graphical display.

11. The method of claim 10 , wherein the audio data is real-time streamed audio data.

12. The method of claim 10 , wherein the remedial action is presented to the specific customer service agent in real-time.

13. The method of claim 10 , wherein the audio data is an interaction between at least the specific customer service agent and at least one customer.

14. The method of claim 10 , wherein at least one of the plurality of utterances consists of more than a single word.

15. A method of script identification in audio data, the method comprising:

receiving audio data;

performing voice activity detection to segment the audio data into a plurality of utterances, wherein each utterance is a segment of speech separated by a segment of non-speech;

filtering the plurality of utterances into a subset of utterances attributed to any customer service agent;

performing speech analytics on the subset of the plurality of utterances to identify at least one event, wherein the event is a detection of at least one keyword in the subset of the plurality of utterances, further wherein the identified event is associated with a requirement that a customer service agent speak at least one of a plurality of script texts in the audio data;

determining the least one script text required to be spoken in the audio data based on the identified at least one event;

receiving at least one script model compilation associated with the at least one script text, wherein each script model compilation includes text from the associated script text and a plurality of acceptable variations for the text of the associated script text;

determining compliance of the plurality of utterances to each script model compilation, wherein determining compliance includes:

transcribing the subset of the plurality of utterances to produce an utterance transcript,

comparing the script model compilation to the utterance transcript to determine a word error rate in the utterance transcript, and

determining the compliance of the utterance transcript based on the word error rate and an error threshold for the script model compilation; and

initiating at least one remedial action if the utterance transcript is non-compliant to present on screen guidance to a customer service agent on a graphical display.

16. The method of claim 15 , wherein the audio data is real-time streamed audio data.

17. The method of claim 15 , wherein the audio data is an interaction between at least a customer service agent and at least one customer.

18. The method of claim 15 , wherein at least one of the plurality of utterances consists of more than a single word.

Assignments (3)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT SYSTEMS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: VERINT SYSTEMS LTD.
To: VERINT SYSTEMS INC.
Reel/Frame 057568/0183 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2020
From: IANNONE, JEFFERY MICHAEL; WEIN, RON; ZIV, OMER
To: VERINT SYSTEMS LTD.
Reel/Frame 052087/0377 →