IP Library Granted Patent US 8,892,438
Granted Patent B2
US 8,892,438 · App. 12/881,665 · Granted Nov 18, 2014

Apparatus and method for analysis of language model changes

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,892,438
App. No.
12/881,665
Granted
Nov 18, 2014
Kind
B2
Abstract

An apparatus, a method, and a machine-readable medium are provided for characterizing differences between two language models. A group of utterances from each of a group of time domains are examined. One of a significant word change or a significant word class change within the plurality of utterances is determined. A first cluster of utterances including a word or a word class corresponding to the one of the significant word change or the significant word class change is generated from the utterances. A second cluster of utterances not including the word or the word class corresponding to the one of the significant word change or the significant word class change is generated from the utterances.

Claims (87)

1. A method comprising:

selecting a plurality of language models;

for each period of a plurality of time periods:

identifying a first utterance and a second utterance received during each time period, wherein the first utterance was recognized using a first language model of the plurality of language models and the second utterance was recognized using a second language model of the plurality of language models;

identifying distinctions between the first utterance and the second utterance for each of the plurality of time periods;

determining when a significant word usage change has occurred within the first language model and the second language model by comparing the distinctions to previously recorded distinctions; and

when the significant word usage change is detected:

identifying a word corresponding to the significant word usage change;

generating, from the utterances, a first cluster of utterances comprising the word;

generating, from the utterances, a second cluster of utterances not comprising the word; and

updating the plurality of language models using the first cluster of utterances and the second cluster of utterances.

2. The method of claim 1 , wherein determining when the significant word usage change has occurred further comprises comparing ones of the utterances from a particular group of speakers associated with one of the plurality of time periods with ones of the utterances from the particular group of speakers from another time period of the plurality of time periods.

3. The method of claim 1 , wherein determining when the significant word usage change has occurred further comprises comparing ones of the utterances from a first particular group of speakers with ones of the utterances from a second particular group of speakers.

4. The method of claim 1 , wherein determining when the significant word usage change has occurred further comprises comparing ones of the utterances from one of the plurality of time periods with ones of the utterances from another of the plurality of time periods.

5. The method of claim 1 , further comprising outputting a list of the utterances in the first cluster and the utterances in the second cluster with analytical information.

6. The method of claim 5 , wherein the analytical information comprises a history of splits that produced each of the clusters.

7. The method of claim 1 , wherein the plurality of time periods comprises two non-overlapping months.

8. The method of claim 1 , further comprising iteratively performing:

examining a group of utterances from the second cluster of utterances, to yield examined utterances;

determining a next significant word change and a next word corresponding to the next significant word change;

generating, from the examined utterances, a new first cluster of utterances comprising the next word; and

generating, from the examined utterances, a new second cluster of utterances not comprising the next word corresponding to the next significant word change.

9. The method of claim 1 , further comprising:

pooling the plurality of utterances from the plurality of time periods, to yield pooled utterances;

assigning each utterance of the pooled utterances to one of a plurality of subpopulations;

generating a language model for each of the plurality of subpopulations;

reassigning each utterance to one of the plurality of subpopulations according to a reassignment criterion;

determining whether any of the plurality of subpopulations fulfill a splitting criterion; and

splitting ones of the plurality of subpopulations that fulfill the splitting criterion, wherein:

examining a subset of the plurality of utterances from each of the time periods, determining one of a significant word change within the plurality of spoken utterances, generating, from the subset of the plurality of utterances, a first cluster of utterances including a word corresponding to the significant word change, and generating, from the subset of the plurality of utterances, a second cluster of utterances not including the word corresponding to the significant word change are performed after pooling, assigning, generating a language model, reassigning, determining whether any of the subpopulations fulfill the splitting criterion, and splitting.

10. The method of claim 9 , wherein examining, determining when a significant word usage change has occurred, generating a first cluster, and generating a second cluster are performed for each of the plurality of subpopulations.

11. The method of claim 9 , wherein:

assigning each utterance of the pooled utterances to one of the plurality of subpopulations further comprises:

assigning each of the utterances of the pooled utterances to one of two subpopulations;

reassigning each utterance of the pooled utterances to one of the plurality of subpopulations according to the reassignment criterion comprises reassigning each utterance of the pooled utterances to one of two subpopulations according to the reassignment criterion; and

splitting ones of the plurality of subpopulations that fulfill the splitting criterion comprises splitting ones of the subpopulations that fulfill the splitting criterion into two subpopulations.

12. The method of claim 9 , further comprising:

iteratively performing, until the language models converge:

generating a language model for each of the subpopulations, and

reassigning each of the plurality of utterances to one of the subpopulations according to the reassignment criterion.

13. The method of claim 9 , wherein the reassignment criterion comprises a subpopulation that maximizes a probability of an utterance occurring.

14. The method of claim 1 , further comprising performing, before performing the acts of claim 1 :

for each of the plurality of time periods, computing a matrix of associational usage scores among all words of a list of frequently occurring words;

computing differences in the associational scores of two of the matrices to produce a difference matrix;

producing a set of clusters of utterances based on similarity in associational usage scores of the difference matrix, to yield a produced set of clusters; and

creating a plurality of word classes based on a result of producing a set of clusters of utterances, wherein:

determining determines a significant word usage class change, and

generating the first cluster comprises generating, from the examined utterances, a cluster of utterances having a word class corresponding to the significant word class change.

15. The method of claim 14 , further comprising:

prioritizing the first cluster of utterances and the second cluster of utterances.

16. A system comprising:

a processor; and

a computer readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

selecting a plurality of language models;

for each time period of a plurality of time periods:

identifying a first utterance and a second utterance received during each time period, wherein the first utterance was recognized using a first language model of the plurality of language models and the second utterance was recognized using a second language model of the plurality of language models;

identifying distinctions between the first utterance and the second utterance for each of the plurality of time periods;

determining when a significant word usage change has occurred within the first language model and the second language model by comparing the distinctions to previously recorded distinctions; and

when the significant word usage change is detected:

identifying a word corresponding to the significant word usage change;

generating, from the utterances, a first cluster of utterances comprising the word;

generating, from the utterances, a second cluster of utterances not comprising the word; and

updating the plurality of language models using the first cluster of utterances and the second cluster of utterances.

17. The system of claim 16 , wherein determining when the significant word usage change has occurred further comprises comparing ones of the utterances from a particular group of speakers of one of the plurality of time periods with ones of the utterances from the particular group of speakers.

18. The system of claim 16 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising:

iteratively performing:

examining a group of utterances from the second cluster of utterances, to yield examined utterances;

determining a next significant word usage change and a next word corresponding to the next significant word usage change;

generating, from the examined utterances, a new first cluster of utterances comprising the next word; and

generating, from the examined utterances, a new second cluster of utterances not comprising the next word.

19. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

selecting a plurality of language models;

for each time period of a plurality of time periods:

identifying a first utterance and a second utterance received during each time period, wherein the first utterance was recognized using a first language model of the plurality of language models and the second utterance was recognized using a second language model of the plurality of language models;

identifying distinctions between the first utterance and the second utterance for each of the plurality of time periods;

determining when a significant word usage change has occurred within the first language model and the second language model by comparing the distinctions to previously recorded distinctions; and

when the significant word usage change is detected:

identifying a word corresponding to the significant word usage change;

generating, from the utterances, a first cluster of utterances comprising the word;

generating, from the utterances, a second cluster of utterances not comprising the word; and

updating the plurality of language models using the first cluster of utterances and the second cluster of utterances.

20. The computer-readable storage device of claim 19 , having additional instructions stored which, when executed by the computing device, result in preliminary operations to be executed by the computing device before the operations of claim 19 , the preliminary operations comprising:

for each time period of a plurality of time periods, computing a matrix of associational usage scores among all words of a list of frequently occurring words;

computing differences in the associational usage scores of two of the matrices to produce a difference matrix;

producing a set of clusters of utterances based on similarity in associational usage scores of the difference matrix, to yield a produced set of clusters; and

creating a plurality of word classes based on a result of producing a set of clusters of utterances, wherein:

determining determines a significant word class change, and generating the first cluster comprises generating, from the examined utterances, a cluster of utterances having a word class corresponding to the significant word class change.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2015
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 035355/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2015
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 035355/0778 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2014
From: GORIN, ALLEN LOUIS; GROTHENDIECK, JOHN; WRIGHT, JEREMY HUNTLEY GREET
To: AT&T CORP.
Reel/Frame 033526/0970 →