IP Library Granted Patent US 9,990,920
Granted Patent B2
US 9,990,920 · App. 15/332,411 · Granted Jun 5, 2018

System and method of automated language model adaptation

Inventors: Ran Achituv (Hod Hasharon, IL); Omer Ziv (Ramat Gan, IL); Ido Shapira (Tel Aviv, IL); Daniel Baum (Modiin, IL)
Assignee: VERINT SYSTEMS LTD.
G10L15/197G10L15/063G10L15/083G10L2015/0635H04M3/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,990,920
App. No.
15/332,411
Granted
Jun 5, 2018
Kind
B2
Abstract

Systems and methods of automated adaptation of a language model for transcription of audio data include obtaining audio data. The audio data is transcribed with a language model to produce a plurality of audio file transcriptions. A quality of the plurality of audio file transcriptions is evaluated. At least one best transcription from a plurality of audio file transcriptions is selected based upon the evaluated quality. Statistics are calculated from the selected at least one best transcription from the plurality of audio file transcriptions. The language model is modified from the calculated statistics.

Claims (59)

1. A method of automated adaptation of a language model for transcription of a call center's audio data, the method comprising:

obtaining transcriptions corresponding to audio data from customer service interactions between one or more customers and one or more customer service agents, wherein each transcription comprises word lattices obtained using the language model;

converting each word lattice for each transcription into a confusion network;

computing an overall quality for each transcription based on the confusion networks corresponding to each transcription;

filtering the transcriptions to retain only transcriptions having an overall quality above a threshold;

calculating language statistics from the retained transcriptions;

modifying the language model using the calculated language statistics; and

using the modified language model for subsequent transcriptions of the call center's audio data.

2. The method of claim 1 , wherein the transcriptions are large vocabulary speech recognition (LVCSR) transcriptions.

3. The method of claim 1 , wherein the converting each word lattice for each transcription into a confusion network comprises:

applying a minimum Bayes risk decoder to each word lattice.

4. An adaptive transcription system for a call center, the adaptive transcription system comprising:

a non-transitory computer readable storage medium storing:

transcriptions corresponding to audio data from customer service interactions between one or more customers and one or more customer service agents, and

a language model that mathematically represents the statistical distribution of words and terms expected in the customer service interactions; and

a processor communicatively coupled to the non-transitory computer readable storage medium, wherein the processor executes computer readable instructions that cause the processor to:

periodically update the language model by:

obtaining the transcriptions corresponding to audio data from customer service interactions during a period, wherein each transcription comprises word lattices obtained using the language model,

converting each word lattice for each transcription into a confusion network,

computing an overall quality for each transcription based on the confusion networks corresponding to each transcription,

filtering the transcriptions to retain only transcriptions having an overall quality above a threshold,

calculating language statistics from the retained transcriptions, and

modifying the language model using the calculated statistics.

5. The adaptive transcription system for a call center according to claim 4 , wherein the modifying the language model using the calculated language statistics comprises:

adapting the language model stored on the non-transient computer readable medium with the modified language model, wherein the adapted language model includes all previous modifications to the language model.

6. The adaptive transcription system for a call center according to claim 4 , wherein the modifying the language model using the calculated language statistics comprises:

replacing the language model stored on the non-transient computer readable medium with the modified language model.

7. A non-transitory computer readable storage medium containing computer readable instructions that when executed by a processor of a computing device cause the computing device to perform a method comprising:

obtaining transcriptions corresponding to audio data from customer service interactions between one or more customers and one or more customer service agents, wherein each transcription comprises word lattices obtained using the language model;

converting each word lattice from each transcription into a confusion network;

computing an overall quality for each transcription based on the confusion networks corresponding to each transcription;

filtering the transcriptions to retain only transcriptions having an overall quality above a threshold;

calculating language statistics from the retained transcriptions;

modifying the language model using the calculated language statistics; and

using the modified language model for subsequent transcriptions of the call center's audio data.

8. The method according to claim 4 , wherein the periodic update of the model occurs after a new product is introduced and the transcriptions correspond to audio data from customer service interactions corresponding to the new product.

9. The method according to claim 1 , wherein the calculating language statistics from the retained transcriptions comprises:

computing occurrences of a word, a word pair, a word triplet, a phrase, or a script in the retained transcriptions, and

determining that the word, the word pair, the word triplet, the phrase, or the script is likely based on the computed number of occurrences.

10. The method according to claim 9 , wherein the calculating language statistics from the retained transcriptions further comprises:

comparing the occurrences of the word, the word pair, the word triplet, the phrase, or the script in the retained transcriptions to a baseline, and

determining that the word, the word pair, the word triplet, the phrase, or the script in the retained transcriptions has increased or decreased from the baseline.

11. The method according to claim 1 , wherein the transcriptions correspond to customer service interactions in a particular language and the modified language model comprises vocabulary and relationships between words corresponding to the particular language.

12. The method according to claim 1 , wherein the transcriptions correspond to customer service interactions in a particular dialect and the modified language model comprises vocabulary and relationships between words corresponding to the particular dialect.

13. The method according to claim 1 , wherein the transcriptions correspond to customer service interactions in a particular industry or field and the modified language model comprises vocabulary and relationships between words corresponding to the particular industry or field.

14. The method according to claim 1 , wherein the transcriptions correspond to customer service interactions are related to a particular service, product, or issue and the modified language model comprises vocabulary and relationships between words corresponding to the particular service, product, or issue.

15. The method according to claim 1 , wherein the obtaining, converting, computing, filtering, calculating, and modifying are performed iteratively to periodically modify the language model used for transcription of the call center's audio data.

16. The method according to claim 1 , wherein the confusion network comprises time segments and one or more word alternatives for each time segment, and wherein each word alternative in a time segment is assigned a probability that it was correctly transcribed.

17. The method according to claim 16 , wherein computing an overall quality for each transcription based on the confusion networks corresponding to each transcription, comprises, for each transcription:

selecting an utterance;

obtaining the confusion network for the utterance;

determining the word alternative with the highest probability in each time segment of the confusion network;

computing a joint probability for adjacent segments using the word alternatives with the highest probability;

averaging the joint probabilities for adjacent segments to determine a quality for the utterance;

repeating the selecting, obtaining, determining, computing, and averaging for other utterances in the transcript; and

determining the overall quality for the transcript as the average of the qualities for each utterance in the transcript.

18. The method according to claim 1 , wherein the filtering the transcriptions to retain only transcriptions having an overall quality above a threshold, comprises:

mapping, using a nonlinear function, the overall quality for each transcription to a quality score between 0 and 1, wherein the mapping enhances the filtering by separating the quality scores for transcriptions near the threshold; and

comparing the quality scores to the threshold.

Assignments (3)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT SYSTEMS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: VERINT SYSTEMS LTD.
To: VERINT SYSTEMS INC.
Reel/Frame 057568/0183 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2016
From: ACHITUV, RAN; ZIV, OMER; SHAPIRA, IDO; BAUM, DANIEL
To: VERINT SYSTEMS LTD.
Reel/Frame 040171/0538 →
Continuity (4)
Continuation 14291895 · May 30, 2014
Provisional Application 61870842 · Aug 28, 2013
Provisional Application 61870843 · Aug 28, 2013
Related Publication 20170098445A1 · Apr 6, 2017