IP Library Granted Patent US 7,962,331
Granted Patent B2
US 7,962,331 · App. 12/255,564 · Granted Jun 14, 2011

System and method for tuning and testing in a speech recognition system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,962,331
App. No.
12/255,564
Granted
Jun 14, 2011
Kind
B2
Abstract

Systems and methods for improving the performance of a speech recognition system. In some embodiments a tuner module and/or a tester module are configured to cooperate with a speech recognition system. The tester and tuner modules can be configured to cooperate with each other. In one embodiment, the tuner module may include a module for playing back a selected portion of a digital data audio file, a module for creating and/or editing a transcript of the selected portion, and/or a module for displaying information associated with a decoding of the selected portion, the decoding generated by a speech recognition engine. In other embodiments, the tester module can include an editor for creating and/or modifying a grammar, a module for receiving a selected portion of a digital audio file and its corresponding transcript, and a scoring module for producing scoring statistics of the decoding based at least in part on the transcript.

Claims (40)

1. A method of tuning a speech recognizer, the method comprising:

playing a selected portion of a digital audio data file with a digital audio player;

creating and/or modifying a digital transcript of the selected audio portion;

displaying information associated with a decode of the selected audio portion on an electronic display, wherein the displayed information includes a menu of selectable noise tags for identifying noise events in the transcript;

receiving an input selected from the displayed menu, the input identifying at least one noise event in the transcript;

modifying, by a computing device, the transcript based on the input; and

utilizing the modified transcript to improve the performance of the speech recognizer.

2. The method of claim 1 , wherein the menu comprises a graphical user interface having elements for allowing selection, input, and command entry related to the selectable noise tags.

3. The method of claim 1 , wherein the information comprises a confidence score.

4. The method of claim 1 , wherein the information comprises an acoustic model score.

5. The method of claim 1 , wherein the at least one noise event is a cough.

6. The method of claim 1 , wherein the at least one noise event is a sneeze.

7. The method of claim 1 , wherein the at least one noise event is a laugh.

8. The method of claim 1 , wherein the at least one noise event is a breath.

9. The method of claim 1 , wherein speech recognizer performance is improved by training new acoustic models.

10. The method of claim 1 , wherein improving the performance of the speech recognizer includes tuning parameters in the speech recognizer that are not acoustic models.

11. The method of claim 1 , wherein improving the performance of the speech recognizer comprises using the identified at least one noise event to train the speech recognizer to interpret or ignore acoustic phenomena characterized as noise.

12. The method of claim 1 , wherein improving the performance of the speech recognizer comprises building a new acoustic model using the identified at least one noise event.

13. The method of claim 1 , further comprising determining, based at least in part on the transcript and the information associated with the decode, a modification of the speech recognizer to improve its performance.

14. The method of claim 13 , wherein the modification comprises modifying a grammar of the speech recognizer.

15. The method of claim 14 , wherein the modification comprises adding a concept, phrase, word, or phoneme to the grammar.

16. The method of claim 13 , wherein the modification comprises modifying a word pronunciation, dictionary, or acoustic model of the speech recognizer.

17. The method of claim 13 , wherein the modification comprises modifying a call flow.

18. The method of claim 13 , further comprising making a modification to the speech recognizer.

19. The method of claim 18 , further comprising iteratively performing the recited steps.

20. A system for facilitating the tuning of a speech recognizer, the system comprising:

a processor;

a memory;

a playback module configured to play a selected portion of a digital audio data file;

a user interface configured to provide a menu of selectable noise tags for identifying noise events in a transcript;

an editor module configured to receive input modifying the transcript or the notes, wherein the input includes noise tags, selected from the menu of selectable noise tags, attaching markers to the transcript; and

a detail viewing module configured to display information associated with a decoding of the selected portion by the speech recognizer, the information including the noise tags identifying noise events in the transcript.

21. The system of claim 20 , further comprising a user interface that includes the menu of selectable noise tags.

22. The system of claim 20 , wherein the user interface comprises a graphical user interface.

23. The system of claim 20 , wherein the information associated with the decoding comprises a grammar associated with the selected portions.

24. The system of claim 23 , wherein the grammar comprises a set of responses expected to occur in the selected portions.

25. The system of claim 24 , wherein the set of responses comprises phrases, words, and/or phonemes.

26. The system of claim 20 , wherein the information associated with the decoding comprises a confidence score.

27. The system of claim 20 , wherein the information associated with the decoding comprises an identification of an acoustic model.

28. The system of claim 20 , wherein the information associated with the decoding comprises phonemes used by the speech recognizer to decode the selected portions.

Assignments (8)
RELEASE OF SECURITY INTEREST Recorded Jul 1, 2025
From: ORIX GROWTH CAPITAL, LLC
To: AI SOFTWARE LLC; TEXTEL CX, INC.; DENIM SOCIAL, LLC; SMARTACTION HOLDINGS, INC.; SMARTACTION LLC
Reel/Frame 071584/0477 →
RELEASE OF SECURITY INTEREST Recorded Jun 20, 2025
From: ESPRESSO CAPITAL LTD.
To: LUMENVOX, LLC
Reel/Frame 071471/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2025
From: LUMENVOX, LLC; LUMENVOX CORPORATION
To: AI SOFTWARE, LLC
Reel/Frame 071362/0692 →
SECURITY INTEREST Recorded Jun 17, 2024
From: AI SOFTWARE LLC; TEXTEL CX, INC.; DENIM SOCIAL, LLC; SMARTACTION HOLDINGS, INC.; SMARTACTION LLC
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 067743/0755 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: ESPRESSO CAPITAL LTD.
To: AI SOFTWARE, LLC; TEXTEL GC, INC.
Reel/Frame 067194/0772 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 3, 2023
From: AI SOFTWARE, LLC; TEXTEL GC, INC.
To: ESPRESSO CAPITAL LTD.
Reel/Frame 064690/0684 →
SECURITY INTEREST Recorded Mar 24, 2021
From: LUMENVOX, LLC
To: ESPRESSO CAPITAL LTD.
Reel/Frame 055705/0320 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2010
From: MILLER, EDWARD S.; BLAKE, JAMES F. II; HEROLD, KEITH C.; BERGMAN, MICHAEL D.; DANIELSON, KYLE N.; AUCKLAND, ALEXANDRA L.
To: LUMENVOX, LLC
Reel/Frame 025170/0903 →