IP Library Granted Patent US 8,473,292
Granted Patent B2
US 8,473,292 · App. 12/552,832 · Granted Jun 25, 2013

System and method for generating user models from transcribed dialogs

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,473,292
App. No.
12/552,832
Granted
Jun 25, 2013
Kind
B2
Abstract

Disclosed herein are systems, computer-implemented methods, and computer-readable storage media for generating personalized user models. The method includes receiving automatic speech recognition (ASR) output of speech interactions with a user, receiving an ASR transcription error model characterizing how ASR transcription errors are made, generating guesses of a true transcription and a user model via an expectation maximization (EM) algorithm based on the error model and the respective ASR output where the guesses will converge to a personalized user model which maximizes the likelihood of the ASR output. The ASR output can be unlabeled. The method can include casting speech interactions as a dynamic Bayesian network with four variables: (s), (u), (r), (m), and encoding relationships between (s), (u), (r), (m) as conditional probability tables. At each dialog turn (r) and (m) are known and (s) and (u) are hidden.

Claims (38)

1. A method comprising:

receiving automatic speech recognition output of a plurality of speech interactions with a user;

receiving an automatic speech recognition transcription error model characterizing how automatic speech recognition transcription errors are made;

generating, via a processor, guesses of a true transcription and a user model via an expectation maximization algorithm, wherein the expectation maximization algorithm is based on the error model and the automatic speech recognition output; and

generating a personalized user model associated with a user voiceprint of the user based on the guesses and the user model.

2. The method of claim 1 , wherein generating the guesses of the true transcription and the user model further comprises iteratively repeating the following steps until a threshold is met:

guessing the true transcription from a current guess of the user model, to yield a current guess of the true transcription; and

guessing the user model based on the current guess of the true transcription.

3. The method of claim 1 , wherein the expectation maximization algorithm estimates conditional probabilities of hidden variables.

4. The method of claim 1 , wherein generating the guesses is further based on a set of manual transcriptions of speech interactions with the user.

5. The method of claim 4 , wherein the set of manual transcriptions is less numerous than the automatic speech recognition output.

6. The method of claim 1 , the method further comprising casting the speech interactions as a dynamical Bayesian network having variables and encoding relationships between the variables as conditional probability tables.

7. The method of claim 6 , wherein the variables comprise a user state, a user action, a speech recognition output, and a dialog system action.

8. The method of claim 1 , the method further comprising generating a personalized automatic speech recognition model based on the personalized user model.

9. The method of claim 1 , the method further comprising recognizing additional speech from the user based on the personalized speech model.

10. The method of claim 9 , the method further comprising iteratively improving the personalized speech model based on the additional speech.

11. The method of claim 1 , wherein the personalized user model applies to one of an individual user, a group of individual users and a population segment of.

12. The method of claim 11 , wherein the personalized user model applied to the segment of similar users is based on an individual user simulation.

13. The method of claim 1 , wherein the automatic speech recognition output is unlabeled.

14. A system comprising:

a processor; and

a computer-readable storage device having instructions stored which, when executed on the processor, perform operations comprising:

receiving automatic speech recognition output of a plurality of speech interactions with a user;

receiving an automatic speech recognition transcription error model characterizing how automatic speech recognition transcription errors are made;

generating, via a processor, guesses of a true transcription and a user model via an expectation maximization algorithm, wherein the expectation maximization algorithm is based on the error model and the respective automatic speech recognition output; and

generating a personalized user model associated with a user voiceprint of the user based on the guesses and the user model.

15. The system of claim 14 , wherein the automatic speech recognition output is unlabeled.

16. The system of claim 14 , wherein generating the guesses of the true transcription and the user model further comprises repeating the following steps until a threshold is met:

guessing the true transcription from a current guess of the user model, to yield a current guess of the true transcription; and

guessing the user model from the current guess of the true transcription.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving a user model personalized for a specific user generated by steps comprising:

receiving automatic speech recognition output of a plurality of speech interactions with the specific user;

receiving an automatic speech recognition transcription error model characterizing how automatic speech recognition transcription errors are made;

generating, via a processor, guesses of a true transcription and a user model via an expectation maximization algorithm, wherein the expectation maximization algorithm is based on the error model and the automatic speech recognition output; and

generating a personalized user model for the specific user based on the guesses and the user model; and

building a personalized dialog system for the specific user based associated with a user voiceprint of the user on the received personalized user model.

18. The computer-readable storage device of claim 17 , wherein the automatic speech recognition output is unlabeled.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2009
From: WILLIAMS, JASON; SYED, UMAR
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 023185/0512 →