IP Library Granted Patent US 7,860,715
Granted Patent B2
US 7,860,715 · App. 11/677,220 · Granted Dec 28, 2010

Method, system and program product for training and use of a voice recognition application

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,860,715
App. No.
11/677,220
Granted
Dec 28, 2010
Kind
B2
Abstract

A system, method and program product for the shortcomings of the prior art are overcome and additional advantages are provided through a system, method and program product for initializing a speech recognition application for a computer. The method comprises recording a variety of sounds associated with a specific text; identifying location of different words as pronounced in different locations of this recorded specific text; and calibrating word location of an input stream based on results of the pre-recorded and identified word locations when attempting to parse words received from spoken sentences of the input stream.

Claims (26)

1. A method for use with a speech recognition application for a computer, comprising:

receiving a plurality of spoken sounds associated with a specific text having words, the specific text comprising at least one predetermined word;

identifying and recording, via at least one computer, locations in the plurality of spoken sounds that correspond to pronunciations of the words in said specific text by recognizing pronunciation of the at least one predetermined word in at least one predetermined position and using recognition of the pronunciation of the at least one predetermined word in the at least one predetermined position to align the plurality of spoken sounds with the specific text; and

storing identification information about pronunciation of the words from the specific text in a variety of sound files to be used later for calibration purposes.

2. The method of claim 1 , wherein said sound files are stored in a memory location.

3. The method of claim 2 , wherein said speech recognition application is utilized in a computing environment having one or more nodes.

4. The method of claim 3 , wherein said nodes comprise one or more computers.

5. The method of claim 3 , wherein said recorded sound files are stored in a memory location accessible to the one or more nodes utilizing said speech recognition application.

6. The method of claim 5 , wherein said sound files are arranged in a database format capable of being arranged in a certain manner.

7. The method of claim 5 , wherein said nodes each have their own independent storage medium for storing said sound files and said nodes are enabled to utilize each others sound files by accessing each of said independent storage medium associated with each of said nodes.

8. The method of claim 7 , wherein said sound files are in MP3 format.

9. The method of claim 7 , wherein said sound files are in WAV format.

10. The method of claim 5 , wherein at least one predetermined word comprises a plurality of synchronization keywords, and wherein during the sound receiving step the plurality of synchronization keywords are established to later identify word location.

11. The method of claim 5 , wherein said sound files are stored in a central storage location accessible by all said nodes.

12. The method of claim 11 , wherein said sound files are in MP3 format.

13. The method of claim 11 , wherein said sound files are in WAV format.

14. The method of claim 11 , wherein said plurality of synchronization keywords are stored in a storage location for later use.

15. The method of claim 14 , wherein said plurality of synchronization keywords allow continuously synchronizing or re-synchronizing of word calibration during identifying the locations of the words and later when used to calibrate word locations.

16. The method of claim 15 , wherein the process of choosing the plurality of synchronization keywords can be pre-selected and altered based on usage and/or other factors.

17. The method of claim 7 , wherein said sound files can be used in establishing peer to peer communication when more than one nodes in said computing environment desire to communicate with one another.

18. A system for speech recognition by a computer system, comprising:

a storage medium for receiving at least one recorded sound file having words, the at least one sound file being associated with a text and comprising synchronization keywords; and

the synchronization keywords for identifying, via at least one computer, locations in a plurality of spoken sounds from the at least one recorded sound file that corresponds to pronunciations of the words in the at least one recorded sound file by recognizing pronunciation of the synchronization keywords in at least one predetermined position and using recognition of the pronunciation of the synchronization keywords in the at least one predetermined position to align the plurality of spoken sounds with the text associated with said at least one recorded sound file.

19. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by at least one computer, perform a method for use with a speech recognition application for a computer, the method comprising:

(a) receiving a plurality of spoken sounds associated with a specific text having words, the specific text comprising at least one predetermined word; and

(b) identifying and recording, by the at least one computer, locations in the plurality of spoken sounds that correspond to pronunciations of the words in said specific text by recognizing pronunciation of the at least one predetermined word in at least one predetermined position and using recognition of the pronunciation of the at least one predetermined word in the at least one predetermined position to align the plurality of spoken sounds with the specific text.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2007
From: ESSENMACHER, MICHAEL D.; MURPHY, THOMAS E., JR.; STEVENS, JEFFREY S.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 018915/0009 →