IP Library Granted Patent US 9,905,224
Granted Patent B2
US 9,905,224 · App. 14/736,277 · Granted Feb 27, 2018

System and method for automatic language model generation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,905,224
App. No.
14/736,277
Granted
Feb 27, 2018
Kind
B2
Abstract

A computer-implemented method of generating a language model. An embodiment of a system and method may include selecting a set of words from a transcription of an audio input, the transcription produced by a current language model. The set of words may be used to obtain a set of content objects. The set of content objects may be used to generate a new language model. The current language model may be replaced by the new language model.

Claims (38)

1. A computer-implemented method of automatically producing an improved transcription, the method comprising:

obtaining, by a processor, an audio input, the audio input comprising a recording of audio signal;

producing, by the processor, a first transcription of the audio input using a current language model;

associating, by the processor, words included in the first transcription with probabilities each probability indicating for an associated word the likelihood that the word is a legitimate word;

selecting, by the processor, a set of words from the first transcription having associated probabilities greater than a threshold;

using, by the processor, the selected set of words to search, in at least one of: the internet and a database, for a set of additional textual objects that include at least one of the selected set of words, wherein the additional textual content objects are selected from the list consisting of: webpages, text posted in a social network, articles published on the internet and textual documents;

using, by the processor, the additional textual objects to train a new language model;

adapting, by the processor, the current language model based on the new language model to produce an improved adapted language model; and

producing, by the processor, a second transcription of the audio input using the adapted language model.

2. The method of claim 1 , comprising:

comparing performance of the current language model with performance of the adapted language model; and

using the adapted language model for decoding audio content if performance of the adapted language model is better than a performance of the current language model.

3. The method of claim 1 , comprising performing unsupervised testing, wherein unsupervised testing includes validating that the adapted language model is better than the current language model without any intervention or effort by a human.

4. A computer-implemented method of automatically producing an improved transcription of an audio input, the method comprising:

using, by a speech engine, a first language model to decode an input audio recording to produce a first transcription;

associating words included in the first transcription with probabilities each probability indicating for an associated word the likelihood that the word is a legitimate word;

selecting a set of words included in the first transcription having associated probabilities greater than a threshold;

using the set of words to search, in at least one of: the internet and a database, for a set of additional textual content objects that include at least one of the selected set of words, wherein the additional textual content objects are selected from the list consisting of: webpages, text posted in a social network, articles published on the internet and textual documents;

using the additional set of textual content objects to train a second language model;

adapting the first language model based on the second language model to produce an improved adapted language model; and

producing a second transcription of the audio input, by a speech engine, using the adapted language model.

5. The method of claim 4 , comprising:

if a performance of the speech engine, when using the second language model, is better than a performance of the speech engine when using the first language model, then configuring the speech engine to use the adapted language model.

6. The method of claim 4 , comprising validating that the adapted language model is better than the first language model without any intervention or effort by a human.

7. The method of claim 4 , comprising including words in the set of words based on a threshold probability.

8. A system for automatically producing an improved transcription, the system comprising:

memory; and

a processor configured to:

obtain an audio input, the audio input comprising a recording of audio signal;

produce a first transcription of the audio input using a current language model;

associate words included in the first transcription with probabilities each probability indicating for an associated word the likelihood that the word is a legitimate word;

select a set of words from the first transcription having associated probabilities greater than a threshold;

use the selected set of words to search, in at least one of: the internet and a database, for a set of additional textual objects that include at least one of the selected set of words, wherein the additional textual content objects are selected from the list consisting of: webpages, text posted in a social network, articles published on the internet and textual documents;

train a new language model based on the additional textual objects;

adapt the current language model based on the new language model to produce an improved adapted language model; and

produce a second transcription of the audio input using the adapted language model.

9. The system of claim 8 , wherein the processor is configured to compare a performance of the current language model with a performance of the adapted language model, and, use the adapted language model for decoding audio content if a performance of the adapted language model is better than a performance of the current language model.

10. The system of claim 8 , wherein the processor is configured to validate that the adapted language model is better than the current language model without any intervention or effort by a human.

Assignments (4)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
PATENT SECURITY AGREEMENT Recorded Dec 6, 2016
From: NICE LTD.; NICE SYSTEMS INC.; AC2 SOLUTIONS, INC.; ACTIMIZE LIMITED; INCONTACT, INC.; NEXIDIA, INC.; NICE SYSTEMS TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 040821/0818 →
CHANGE OF NAME Recorded Oct 18, 2016
From: NICE-SYSTEMS LTD.
To: NICE LTD.
Reel/Frame 040387/0527 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2015
From: NISSAN, MAOR
To: NICE-SYSTEMS LTD.
Reel/Frame 036242/0256 →