IP Library Granted Patent US 7,092,884
Granted Patent B2
US 7,092,884 · App. 10/086,395 · Granted Aug 15, 2006

Method of nonvisual enrollment for speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,092,884
App. No.
10/086,395
Granted
Aug 15, 2006
Kind
B2
Abstract

In a speech recognition system, a method of nonvisual enrollment comprising playing an audio representation of an enrollment script. As the enrollment is playing, shadowed speech from a user can be received, wherein the shadowed speech can lag the enrollment script. The received shadowed speech can be recorded for enrolling the user into the speech recognition system.

Claims (48)

1. In a speech recognition system, a method of nonvisual enrollment comprising:

playing an audio representation of an enrollment script;

as said enrollment script is playing, receiving a speech sample comprising at least a predetermined minimum amount of shadowed speech from a user wherein said shadowed speech lags the enrollment script;

receiving additional shadowed speech;

selectively replacing a portion of said speech sample with a portion of said additional shadowed speech; and

recording said received shadowed speech for enrolling the user into the speech recognition system.

2. The method of claim 1 , further comprising:

enrolling the user in the speech recognition system by constructing acoustic models based upon the enrollment script and said received shadowed speech.

3. The method of claim 1 , wherein said playing step comprises:

playing a recording of a human voice dictating the enrollment script.

4. The method of claim 1 , wherein said playing step comprises:

playing the enrollment script using a text-to-speech system.

5. The method of claim 1 , further comprising:

pausing said playing of the enrollment script responsive to a user input.

6. The method of claim 5 , further comprising:

resuming said playing of the enrollment script responsive to a user input.

7. The method of claim 1 , further comprising:

monitoring said received shadowed speech and said playing of said enrollment script; and

selectively altering the playback speed of the enrollment script according to said monitoring step.

8. The method of claim 1 , said receiving shadowed speech step furthcr comprising;

receiving a speech sample comprising more than a predetermined minimum amount of shadowed user speech; and

selectively excluding a portion of said speech sample from said enrollment step.

9. The method of claim 1 , wherein said receiving shadowed speech step comprises:

receiving shadowed speech substantially simultaneously with said playing of the enrollment script.

10. A machine-readable storage, having stored thereon a computer program having a plurality of code sections executable by a machine for causing the machine to perform the steps of:

playing an audio representation of an enrollment script;

as said enrollment script is playing, receiving a speech sample comprising at least a predetermined minimum amount of shadowed speech from a user wherein said shadowed speech lags the enrollment script;

receiving additional shadowed speech;

selectively replacing a portion of said speech sample with a portion of said additional shadowed speech; and

recording said received shadowed speech for enrolling the user into the speech recognition system.

11. The machine-readable storage of claim 10 , further comprising:

enrolling the user in the speech recognition system by constructing acoustic models based upon the enrollment script and said received shadowed speech.

12. The machine-readable storage of claim 10 , wherein said playing step comprises:

playing a recording of a human voice dictating the enrollment script.

13. The machine-readable storage of claim 10 , wherein said playing step comprises:

playing the enrollment script using a text-to-speech system.

14. The machine-readable storage of claim 10 , further comprising:

pausing said playing of the enrollment script responsive to a user input.

15. The machine-readable storage of claim 14 , further comprising:

resuming said playing of the enrollment script responsive to a user input.

16. The machine-readable storage of claim 10 , further comprising;

monitoring said received shadowed speech and said playing of said enrollment script; and

selectively altering the playback speed of the enrollment script according to said monitoring step.

17. The machine-readable storage of claim 11 , said receiving shadowed speech step further comprising:

receiving a speech sample comprising more than a predetermined minimum amount of shadowed user speech; and

selectively excluding a portion of said speech sample from said enrollment step.

18. The machine-readable storage of claim 10 , wherein said receiving shadowed speech step comprises:

receiving shadowed speech substantially simultaneously with said playing of the enrollment script.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022354/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2002
From: LEWIS, JAMES R.; POLKOSKY, MELANIE D.; SADOWSKI, JR., WALLACE J.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 012644/0699 →