IP Library Granted Patent US 9,454,966
Granted Patent B2
US 9,454,966 · App. 13/926,552 · Granted Sep 27, 2016

System and method for generating user models from transcribed dialogs

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,454,966
App. No.
13/926,552
Granted
Sep 27, 2016
Kind
B2
Abstract

Disclosed herein are systems, computer-implemented methods, and computer-readable storage media for generating personalized user models. The method includes receiving automatic speech recognition (ASR) output of speech interactions with a user, receiving an ASR transcription error model characterizing how ASR transcription errors are made, generating guesses of a true transcription and a user model via an expectation maximization (EM) algorithm based on the error model and the respective ASR output where the guesses will converge to a personalized user model which maximizes the likelihood of the ASR output. The ASR output can be unlabeled. The method can include casting speech interactions as a dynamic Bayesian network with four variables: (s), (u), (r), (m), and encoding relationships between (s), (u), (r), (m) as conditional probability tables. At each dialog turn (r) and (m) are known and (s) and (u) are hidden.

Claims (40)

1. A method comprising:

receiving speech recognition output associated with speech from a user;

receiving an error model characterizing how automatic speech recognition transcription errors are made;

generating, via a processor, guesses of a true transcription based on an error model-based algorithm, wherein the error model-based algorithm uses the error model and the speech recognition output; and

generating, based on the guesses of the true transcription, a personalized user model associated with a user voiceprint of the user, wherein the generating of the personalized user model comprises iteratively guessing the true transcription until a threshold is met.

2. The method of claim 1 , wherein generating of the guesses further comprises repeating, until a threshold is met, steps comprising:

guessing the true transcription from a current guess of the personalized user model, to yield a current guess of the true transcription; and

guessing the personalized user model based on the current guess of the true transcription.

3. The method of claim 1 , wherein the error model-based algorithm estimates conditional probabilities of hidden variables.

4. The method of claim 1 , wherein generating of the guesses is further based on a set of manual transcriptions of speech interactions with the user.

5. The method of claim 1 , wherein the personalized user model comprises a Bayesian network.

6. The method of claim 1 , further comprising generating a personalized automatic speech recognition model based on the personalized user model.

7. The method of claim 1 , further comprising recognizing additional speech from the user using the personalized user model.

8. The method of claim 7 , further comprising iteratively improving the personalized user model based on the additional speech.

9. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving speech recognition output associated with speech from a user;

receiving an error model characterizing how automatic speech recognition transcription errors are made;

generating guesses of a true transcription based on an error model-based algorithm, wherein the error model-based algorithm uses the error model and the speech recognition output; and

generating, based on the guesses of the true transcription, a personalized user model associated with a user voiceprint of the user, wherein the generating of the personalized user model comprises iteratively guessing the true transcription until a threshold is met.

10. The system of claim 9 , wherein generating of the guesses further comprises repeating, until a threshold is met, steps comprising:

guessing the true transcription from a current guess of the personalized user model, to yield a current guess of the true transcription; and

guessing the personalized user model based on the current guess of the true transcription.

11. The system of claim 9 , wherein the error model-based algorithm estimates conditional probabilities of hidden variables.

12. The system of claim 9 , wherein generating of the guesses is further based on a set of manual transcriptions of speech interactions with the user.

13. The system of claim 9 , wherein the personalized user model comprises a Bayesian network.

14. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising generating a personalized automatic speech recognition model based on the personalized user model.

15. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising recognizing additional speech from the user using the personalized user model.

16. The system of claim 15 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising iteratively improving the personalized user model based on the additional speech.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, result in the computing device performing operations comprising:

receiving speech recognition output associated with speech from a user;

receiving an error model characterizing how automatic speech recognition transcription errors are made;

generating guesses of a true transcription based on an error model-based algorithm, wherein the error model-based algorithm uses the error model and the speech recognition output; and

generating, based on the guesses of the true transcription, a personalized user model associated with a user voiceprint of the user, wherein the generating of the personalized user model comprises iteratively guessing the true transcription until a threshold is met.

18. The computer-readable storage device of claim 17 , wherein generating of the guesses further comprises repeating, until a threshold is met, steps comprising:

guessing the true transcription from a current guess of the personalized user model, to yield a current guess of the true transcription; and

guessing the personalized user model based on the current guess of the true transcription.

19. The computer-readable storage device of claim 17 , wherein the error model-based algorithm estimates conditional probabilities of hidden variables.

20. The computer-readable storage device of claim 17 , wherein generating of the guesses is further based on a set of manual transcriptions of speech interactions with the user.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2013
From: WILLIAMS, JASON; SYED, UMAR
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 030690/0101 →