IP Library Granted Patent US 8,082,148
Granted Patent B2
US 8,082,148 · App. 12/109,204 · Granted Dec 20, 2011

Testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,082,148
App. No.
12/109,204
Granted
Dec 20, 2011
Kind
B2
Abstract

Methods, systems, and products for testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise that include: receiving recorded background noise for each of the plurality of operating environments; generating a test speech utterance for recognition by a speech recognition engine using a grammar; mixing the test speech utterance with each recorded background noise, resulting in a plurality of mixed test speech utterances, each mixed test speech utterance having different background noise; performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances; and evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise.

Claims (56)

1. A computer-implemented method of testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise, the method comprising:

receiving recorded background noise for each of the plurality of operating environments;

generating a test speech utterance for recognition by a speech recognition engine using a grammar;

mixing the test speech utterance with each recorded background noise, resulting in a plurality of mixed test speech utterances, each mixed test speech utterance having different background noise;

performing, for each of the mixed test speech utterances, speech recognition using a device having an automated speech recognition engine, the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise.

2. The method of claim 1 wherein:

performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances further comprises repeatedly performing, for one of the mixed test speech utterances, the speech recognition using the grammar and the one of the mixed test speech utterances, resulting in a plurality of speech recognition results for the one of the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise further comprises calculating, in dependence upon the plurality of speech recognition results for the one of the mixed test speech utterances, a reliability indicator for the grammar when recognizing speech having the recorded background noise included in the one of the mixed test speech utterances.

3. The method of claim 1 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises sampling the test speech utterance from speech of a person.

4. The method of claim 1 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises synthesizing the test speech utterance using a text-to-speech engine.

5. The method of claim 1 wherein testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise further comprises testing a plurality of grammars for use in speech recognition for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction that include a voice mode and one or more non-voice modes, the method further comprising:

identifying, by the multimodal application, a current background noise for a current operating environment in which the multimodal device operates; and

altering, by the multimodal application, flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates.

6. The method of claim 5 wherein altering, by the multimodal application, flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates further comprises:

selecting, by the multimodal application, one of the plurality of grammars tested in dependence upon the current background noise and the evaluation of the speech recognition reliability of the plurality of grammars using the recorded background noises;

receiving, by the multimodal application, a voice utterance from a user; and

performing, by the multimodal application, speech recognition in dependence upon the selected grammar and the voice utterance.

7. A system for testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise, the system comprising:

one or more computer processors; and

computer memory operatively coupled to the one or more computer processors, the computer memory having disposed within it computer program instructions that, when executed by the one or more computer processors, perform acts of:

receiving recorded background noise for each of the plurality of operating environments;

generating a test speech utterance for recognition by a speech recognition engine using a grammar;

mixing the test speech utterance with each recorded background noise, resulting in a plurality of mixed test speech utterances, each mixed test speech utterance having different background noise;

performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise.

8. The system of claim 7 wherein:

performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances further comprises repeatedly performing, for one of the mixed test speech utterances, the speech recognition using the grammar and one of the the mixed test speech utterances, resulting in a plurality of speech recognition results for the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise further comprises calculating, in dependence upon the plurality of speech recognition results for the one of the mixed test speech utterances, a reliability indicator for the grammar when recognizing speech having the recorded background noise included in the one of the mixed test speech utterances.

9. The system of claim 7 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises sampling the test speech utterance from speech of a person.

10. The system of claim 7 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises synthesizing the test speech utterance using a text-to-speech engine.

11. The system of claim 7 wherein testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise further comprises testing a plurality of grammars for use in speech recognition for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction that include a voice mode and one or more non-voice modes, and wherein the computer memory has disposed within it computer program instructions capable of:

identifying a current background noise for a current operating environment in which the multimodal device operates; and

altering flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates.

12. The system of claim 11 wherein altering flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates further comprises:

selecting one of the plurality of grammars tested in dependence upon the current background noise and the evaluation of the speech recognition reliability of the plurality of grammars using the recorded background noises;

receiving a voice utterance from a user; and

performing speech recognition in dependence upon the selected grammar and the voice utterance.

13. A computer program product comprising at least one non-transitory computer readable medium encoded with a plurality of instructions for testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise, the instructions, when executed, perform acts of:

receiving recorded background noise for each of the plurality of operating environments;

generating a test speech utterance for recognition by a speech recognition engine using a grammar;

mixing the test speech utterance with each recorded background noise, resulting in a plurality of mixed test speech utterances, each mixed test speech utterance having different background noise;

performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise.

14. The computer program product of claim 13 wherein:

performing, for each of the mixed test speech utterances, speech recognition using the grammar and the mixed test speech utterance, resulting in speech recognition results for each of the mixed test speech utterances further comprises repeatedly performing, for one of the mixed test speech utterances, the speech recognition using the grammar and the one of the mixed test speech utterances, resulting in a plurality of speech recognition results for one of the mixed test speech utterances; and

evaluating, for each recorded background noise, speech recognition reliability of the grammar in dependence upon the speech recognition results for the mixed test speech utterance having that recorded background noise further comprises calculating, in dependence upon the plurality of speech recognition results for the one of the mixed test speech utterances, a reliability indicator for the grammar when recognizing speech having the recorded background noise included in the one of the mixed test speech utterances.

15. The computer program product of claim 13 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises sampling the test speech utterance from speech of a person.

16. The computer program product of claim 13 wherein generating a test speech utterance for recognition by a speech recognition engine using a grammar further comprises synthesizing the test speech utterance using a text-to-speech engine.

17. The computer program product of claim 13 wherein testing a grammar used in speech recognition for reliability in a plurality of operating environments having different background noise further comprises testing a plurality of grammars for use in speech recognition for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction that include a voice mode and one or more non-voice modes, the method further comprising:

identifying, by the multimodal application, a current background noise for a current operating environment in which the multimodal device operates; and

altering, by the multimodal application, flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates.

18. The computer program product of claim 17 wherein altering, by the multimodal application, flow of execution for the multimodal application in dependence upon the identified current background noise for the current operating environment in which the multimodal application operates further comprises:

selecting, by the multimodal application, one of the plurality of grammars tested in dependence upon the current background noise and the evaluation of the speech recognition reliability of the plurality of grammars using the recorded background noises;

receiving, by the multimodal application, a voice utterance from a user; and

performing, by the multimodal application, speech recognition in dependence upon the selected grammar and the voice utterance.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2008
From: AGAPI, CIPRIAN; BODIN, WILLIAM K; CROSS, CHARLES W, JR; MIRT, MICHAEL H
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 021180/0543 →