IP Library Granted Patent US 9,484,024
Granted Patent B2
US 9,484,024 · App. 14/666,548 · Granted Nov 1, 2016

System and method for handling repeat queries due to wrong ASR output by modifying an acoustic, a language and a semantic model

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,484,024
App. No.
14/666,548
Granted
Nov 1, 2016
Kind
B2
Abstract

Disclosed herein are systems, computer-implemented methods, and computer-readable storage media for handling expected repeat speech queries or other inputs. The method causes a computing device to detect a misrecognized speech query from a user, determine a tendency of the user to repeat speech queries based on previous user interactions, and adapt a speech recognition model based on the determined tendency before an expected repeat speech query. The method can further include recognizing the expected repeat speech query from the user based on the adapted speech recognition model. Adapting the speech recognition model can include modifying an acoustic model, a language model, and a semantic model. Adapting the speech recognition model can also include preparing a personalized search speech recognition model for the expected repeat query based on usage history and entries in a recognition lattice. The method can include retaining unmodified speech recognition models with adapted speech recognition models.

Claims (46)

1. A method comprising:

identifying, based on past interactions with a user participating in a dialog with a speech dialog system, an adaptation schema which, when applied to a speech recognition model, increases a likelihood the speech recognition model will recognize misrecognized speech from the user relative to an unadapted speech recognition model;

determining, via a processor configured to perform speech recognition, that the user has previously repeated speech inputs based on interactions with the user prior to initiating the dialog, to yield a determination; and

adapting, via the processor and based on the determination, the speech recognition model using the adaptation schema before an expected repeat speech input, wherein adapting the speech recognition model further comprises modifying an acoustic model, a language model, and a semantic model.

2. The method of claim 1 , further comprising recognizing the expected repeat speech input from the user based on an adapted speech recognition model.

3. The method of claim 1 , wherein adapting the speech recognition model further comprises preparing a personalized search speech recognition model for the expected repeat speech input based on a usage history of the user and entries in a recognition lattice.

4. The method of claim 1 , further comprising retaining an unmodified speech recognition model in parallel with an adapted speech recognition model.

5. The method of claim 1 , further comprising:

recognizing a repeat input query with an unmodified speech recognition model and with an adapted speech recognition model;

determining a recognition certainty for the unmodified speech recognition model and the adapted speech recognition model; and

basing further interaction with the user on the recognition certainty.

6. The method of claim 1 , further comprising:

determining likely speech characteristics of the expected repeat speech input; and

tailoring an adapted speech recognition model to the likely speech characteristics of the expected repeat speech input.

7. The method of claim 1 , further comprising recording user behavior in a speech input history.

8. A system comprising:

a processor configured to perform speech recognition; and

a computer-readable storage medium having instruction stored which, when executed by the processor, cause the processor to perform operations comprising:

identifying, based on past interactions with a user participating in a dialog with a speech dialog system, an adaptation schema which, when applied to a speech recognition model, increases a likelihood the speech recognition model will recognize misrecognized speech from the user relative to an unadapted speech recognition model;

determining that the user has previously repeated speech inputs based on interactions with the user prior to initiating the dialog, to yield a determination; and

adapting, based on the determination, the speech recognition model using the adaptation schema before an expected repeat speech input, wherein adapting the speech recognition model further comprises modifying an acoustic model, a language model, and a semantic model.

9. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising recognizing the expected repeat speech input from the user based on an adapted speech recognition model.

10. The system of claim 8 , wherein adapting the speech recognition model further comprises preparing a personalized search speech recognition model for the expected repeat speech input based on a usage history of the user and entries in a recognition lattice.

11. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising retaining an unmodified speech recognition model in parallel with an adapted speech recognition model.

12. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

recognizing a repeat input query with an unmodified speech recognition model and with an adapted speech recognition model;

determining a recognition certainty for the unmodified speech recognition model and the adapted speech recognition model; and

basing further interaction with the user on the recognition certainty.

13. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

determining likely speech characteristics of the expected repeat speech input; and

tailoring an adapted speech recognition model to the likely speech characteristics of the expected repeat speech input.

14. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising recording user behavior in a speech input history.

15. A computer-readable storage device having instructions stored which, when executed by a computing device configured to perform speech recognition, cause the computing device to perform operations comprising:

identifying, based on past interactions with a user participating in a dialog with a speech dialog system, an adaptation schema which, when applied to a speech recognition model, increases a likelihood the speech recognition model will recognize misrecognized speech from the user relative to an unadapted speech recognition model;

determining that the user has previously repeated speech inputs based on interactions with the user prior to initiating the dialog, to yield a determination; and

adapting, based on the determination, the speech recognition model using the adaptation schema before an expected repeat speech input, wherein adapting the speech recognition model further comprises modifying an acoustic model, a language model, and a semantic model.

16. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising recognizing the expected repeat speech input from the user based on an adapted speech recognition model.

17. The computer-readable storage device of claim 15 , wherein adapting the speech recognition model further comprises preparing a personalized search speech recognition model for the expected repeat speech input based on a usage history of the user and entries in a recognition lattice.

18. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising retaining an unmodified speech recognition model in parallel with an adapted speech recognition model.

19. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising:

recognizing a repeat input query with an unmodified speech recognition model and with an adapted speech recognition model;

determining a recognition certainty for the unmodified speech recognition model and the adapted speech recognition model; and

basing further interaction with the user on the recognition certainty.

20. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising:

determining likely speech characteristics of the expected repeat speech input; and

tailoring an adapted speech recognition model to the likely speech characteristics of the expected repeat speech input.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2015
From: LJOLJE, ANDREJ; CASEIRO, DIAMANTINO ANTONIO
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 035243/0932 →