IP Library Granted Patent US 7,203,652
Granted Patent B1
US 7,203,652 · App. 10/081,177 · Granted Apr 10, 2007

Method and system for improving robustness in a speech system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,203,652
App. No.
10/081,177
Granted
Apr 10, 2007
Kind
B1
Abstract

The present invention introduces a system and method for improved robustness in a speech system. In one embodiment, a method comprises receiving an utterance from an intended talker at a speech recognition system; computing a speaker verification score with a voice characteristic model associated and with the utterance; computing a speech recognition score associated with the utterance; and selecting a best hypothesis associated with the utterance and based on both the speaker verification score and the speech recognition score.

Claims (70)

1. A method comprising:

receiving a first utterance from a speaker at an integrated speech and speaker recognition system;

generating a voice characteristic model for the speaker;

receiving a second utterance from the speaker at the speaker recognition system;

processing a portion of speech associated with the second utterance, wherein processing comprises,

computing a speaker verification score based on the voice characteristic model associated with the portion of speech,

computing a speech recognition score associated with the portion of speech, and

generating a combined score by combining the speaker verification score and the speech recognition score; and

selecting a best hypothesis from a plurality of hypotheses representing automatic speech recognition results of the second utterance,

based upon the combined score.

2. The method of claim 1 , wherein the portion of speech includes a word, a sentence, a syllable, or a frame.

3. The method of claim 1 , wherein said processing further comprises altering a search path in a Viterbi search used by a speech recognizer.

4. The method of claim 1 , further comprising using hotword speech recognition to identify the speaker.

5. The method of claim 1 , wherein the voice characteristic model includes a voice print, a personal profile and linguistic characteristics.

6. A system comprising:

a speech system; and

a speech input device connected to the speech system; wherein the speech system comprises,

a voice server, wherein the server includes an integrated speech and speaker recognizer that,

receives a first utterance from a speaker via the speech input device;

creates a voice characteristic model for the speaker;

receives a second utterance from the speaker via the speech input device;

processes a portion of speech associated with the second utterance, wherein the processor

computes a speaker verification score based on the voice characteristic model associated with the portion of speech,

computes a speech recognition score associated with the portion of speech, and

generates a combined score by combining the speaker verification score and the speech recognition score; and

selects a best hypothesis from a plurality of hypotheses representing automatic speech recognition results of the second utterance, based upon the combined score.

7. The system of claim 6 , wherein the speech input device comprises a cellular telephone, an analog telephone, a digital telephone, and a voice over internet protocol device.

8. The system of claim 6 , wherein the portion of speech includes a word, a sentence, a syllable, or a frame.

9. The system of claim 6 , wherein the server is further configured to alter a search path in a Viterbi search used by a speech recognizer.

10. An integrated speech and speaker recognition system comprising:

means for receiving a first utterance from a speaker;

means for generating a voice characteristic model for the speaker;

means for receiving a second utterance from the speaker at the speaker recognition system;

means for processing a portion of speech associated with the second utterance, wherein said means for processing comprises,

means for computing a speaker verification score based on the voice characteristic model associated with the portion of speech,

means for computing a speech recognition score associated with the portion of speech, and

means for generating a combined score by combining the speaker verification score and the speech recognition score; and

means for selecting a best hypothesis from a plurality of hypotheses representing automatic speech recognition results of the second utterance, based upon the combined score.

11. The system of claim 10 , wherein the portion of speech includes a word, a sentence, a syllable, or a frame.

12. The system of claim 10 , wherein the means for processing further comprises means for altering a search path in a Viterbi search used by a speech recognizer on the second utterance.

13. The system of claim 10 , further comprising means for using hotword speech recognition to identify the speaker.

14. The system of claim 10 , wherein the voice characteristic model includes a voice print, a personal profile and linguistic characteristics.

15. A machine-readable medium having stored thereon a plurality of instructions which, when executed by a machine, cause said machine to perform a process comprising:

receiving a first utterance from a speaker at an integrated speech and speaker recognition system;

generating a voice characteristic model for the speaker;

receiving a second utterance from the speaker at the speaker recognition system;

processing a portion of speech associated with the second utterance, wherein processing comprises,

computing a speaker verification score based on the voice characteristic model associated with the portion of speech,

computing a speech recognition score associated with the portion of speech, and

generating a combined score by combining the speaker verification score and the speech recognition score; and

selecting a best hypothesis from a plurality of hypotheses representing automatic speech recognition results of the second utterance, based upon the combined score.

16. The machine-readable medium of claim 15 wherein the portion of speech includes a word, a sentence, a syllable, or a frame.

17. The machine-readable medium of claim 15 , having stored thereon additional instructions when processing a portion of speech, said additional instructions when executed by a machine, cause said machine to perform altering a search path in a Viterbi search used by a speech recognizer.

18. The machine-readable medium of claim 15 , having stored thereon additional instructions which, when executed by the machine while identifying a speaker, cause said machine to use hotword speech recognition to identify the speaker.

19. The machine-readable medium of claim 15 , wherein the voice characteristic model includes a voice print, a personal profile and linguistic characteristics.

20. A method comprising:

receiving an utterance from a speaker at a speech recognition system;

computing a speaker verification score based on a voice characteristic model associated and with the utterance;

computing a speech recognition score associated with the utterance; and

selecting a best hypothesis from a plurality of hypotheses representing automatic speech recognition results of the utterance, based on both the speaker verification score and the speech recognition score.

21. The method of claim 20 , wherein the voice characteristic model is obtained from a voice model database.

22. The method of claim 20 , wherein the voice characteristic model is obtained from a first portion of the utterance.

23. A speech recognition system comprising:

a speaker verifier;

a speech recognizer connected to the speaker verifier; and

an input device connected to the speaker verifier and speech recognizer,

wherein the input device receives an utterance from a speaker; and

wherein the speech recognizer generates a recognition score associated with the utterance and generates a plurality of hypotheses representing automatic speech recognition results of the utterance, the speaker verifier generates a speaker verification score associated with the utterance; and the recognition score is combined with the verification score to select a best hypothesis of the plurality of hypotheses.

24. The speech recognition system of claim 23 , wherein the speech recognizer and speaker verifier are software entities residing on a speech server, and wherein the speech server comprises a processor, a bus connected to the processor, and memory connected to the bus that stores the software entities.

25. The speech recognition system of claim 24 , further comprising a database connected to the speech server, wherein the database stores a voice characteristic model of the speaker.

Assignments (4)
PATENT RELEASE (REEL:017435/FRAME:0199) Recorded May 20, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
To: NUANCE COMMUNICATIONS, INC., AS GRANTOR; ART ADVANCED RECOGNITION TECHNOLOGIES, INC., A DELAWARE CORPORATION, AS GRANTOR; SPEECHWORKS INTERNATIONAL, INC., A DELAWARE CORPORATION, AS GRANTOR; TELELOGUE, INC., A DELAWARE CORPORATION, AS GRANTOR; DSP, INC., D/B/A DIAMOND EQUIPMENT, A MAINE CORPORATON, AS GRANTOR; SCANSOFT, INC., A DELAWARE CORPORATION, AS GRANTOR; DICTAPHONE CORPORATION, A DELAWARE CORPORATION, AS GRANTOR
Reel/Frame 038770/0824 →
PATENT RELEASE (REEL:018160/FRAME:0909) Recorded May 20, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
To: NUANCE COMMUNICATIONS, INC., AS GRANTOR; ART ADVANCED RECOGNITION TECHNOLOGIES, INC., A DELAWARE CORPORATION, AS GRANTOR; SPEECHWORKS INTERNATIONAL, INC., A DELAWARE CORPORATION, AS GRANTOR; TELELOGUE, INC., A DELAWARE CORPORATION, AS GRANTOR; DSP, INC., D/B/A DIAMOND EQUIPMENT, A MAINE CORPORATON, AS GRANTOR; HUMAN CAPITAL RESOURCES, INC., A DELAWARE CORPORATION, AS GRANTOR; INSTITIT KATALIZA IMENI G.K. BORESKOVA SIBIRSKOGO OTDELENIA ROSSIISKOI AKADEMII NAUK, AS GRANTOR; NOKIA CORPORATION, AS GRANTOR; MITSUBISH DENKI KABUSHIKI KAISHA, AS GRANTOR; STRYKER LEIBINGER GMBH & CO., KG, AS GRANTOR; NORTHROP GRUMMAN CORPORATION, A DELAWARE CORPORATION, AS GRANTOR; SCANSOFT, INC., A DELAWARE CORPORATION, AS GRANTOR; DICTAPHONE CORPORATION, A DELAWARE CORPORATION, AS GRANTOR
Reel/Frame 038770/0869 →
SECURITY AGREEMENT Recorded Apr 7, 2006
From: NUANCE COMMUNICATIONS, INC.
To: USB AG, STAMFORD BRANCH
Reel/Frame 017435/0199 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2002
From: HECK, LARRY PAUL
To: NUANCE COMMUNICATIONS
Reel/Frame 012638/0296 →