IP Library Granted Patent US 8,306,815
Granted Patent B2
US 8,306,815 · App. 11/951,904 · Granted Nov 6, 2012

Speech dialog control based on signal pre-processing

Assignee: Nuance Communications, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,306,815
App. No.
11/951,904
Granted
Nov 6, 2012
Kind
B2
Abstract

A speech dialog system interfaces a user to a computer. The system includes a signal pre-processor that processes a speech input to generate an enhanced signal and an analysis signal. A speech recognition unit may generate a recognition result based on the enhanced signal. A control unit may manage an output unit or an external device based on the information within the analysis signal.

Claims (63)

1. A speech dialog system, comprising:

a processor;

a signal pre-processor unit configured to process a speech input signal and generate an enhanced speech signal and an analysis signal, where the analysis signal comprises information related to one or more non-semantic characteristics of the speech input signal, the signal pre-processor unit outputting the analysis signal for speech dialog control;

a speech recognition unit configured to receive the enhanced speech signal and generate a recognition result containing one or more words spoken in the speech input signal based on the enhanced speech signal;

a speech output unit configured to output a synthesized speech output in response to the recognition result; and

a speech dialog control unit configured to receive the analysis signal and the recognition result, the speech dialog control unit configured to control the speech output unit based upon the received analysis signal, and the speech dialog control unit also configured to control the signal pre-processor unit based upon the on the recognition result.

2. The speech dialog system of claim 1 , where the information related to one or more non-semantic characteristics of the speech input signal comprises information related to one or more of:

a noise component of the speech input signal;

an echo component of the speech input signal;

a location of a source of the speech input signal; a volume level of the speech input signal;

a pitch of the speech input signal; or

a stationarity of the speech input signal.

3. The speech dialog system of claim 1 , where the speech input signal comprises a representation of one or more words spoken by a user; and

where the information of the analysis signal is unrelated to a meaning or identity of the one or more words spoken by the user.

4. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine a noise component of the speech input signal; and

where the analysis signal comprises information related to the noise component of the speech input signal.

5. The speech dialog system of claim 4 , where the speech dialog control unit is configured to increase a volume of the speech output unit when the noise component of the speech input signal is above a predetermined threshold.

6. The speech dialog system of claim 4 , where the speech dialog control unit is configured to generate an output message to instruct a user to neutralize a noise source when the noise component of the speech input signal is above a predetermined threshold.

7. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine an echo component of the speech input signal; and

where the analysis signal comprises information related to the echo component of the speech input signal.

8. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine a location of a source of the speech input signal; and

where the analysis signal comprises information related to the location of the source of the speech input signal.

9. The speech dialog system of claim 8 , where the speech dialog control unit is configured to customize content for the speech output unit or control an external device based on the location of the source of the speech input signal.

10. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine a volume level of the speech input signal; and

where the analysis signal comprises information related to the volume level of the speech input signal.

11. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine a pitch of the speech input signal; and

where the analysis signal comprises information related to the pitch of the speech input signal.

12. The speech dialog system of claim 11 , where the speech dialog control unit is configured to analyze the information related to the pitch of the speech input signal to determine an age or gender of a user that provided the speech input signal; and

where the speech dialog control unit is configured to customize content for the speech output unit or control an external device based on the age or gender of the user.

13. The speech dialog system of claim 1 , where the signal pre-processor unit is configured to determine a stationarity of the speech input signal; and

where the analysis signal comprises information related to the stationarity of the speech input signal.

14. The speech dialog system of claim 1 , where the speech dialog control unit is configured to use the information of the analysis signal to control the signal pre-processor unit.

15. The speech dialog system of claim 1 , where the signal pre-processor unit comprises a noise reduction filter or an echo compensation filter; and

where the speech dialog control unit is configured to adjust one or more parameters of the noise reduction filter or the echo compensation filter based on the information of the analysis signal.

16. The speech dialog system of claim 1 , where the speech dialog control unit is configured to use the information of the analysis signal to control the speech recognition unit.

17. The speech dialog system of claim 1 , where the speech recognition unit comprises a plurality of available code books for speech recognition; and

where the speech dialog control unit is configured to select one of the plurality of code books to be used by the speech recognition unit based on the information of the analysis signal.

18. The speech dialog system of claim 1 , further comprising:

one or more microphones configured to detect the speech input signal, where the one or more microphones comprise one or more directional microphones; and

where the signal pre-processor unit comprises a beam-former configured to determine a location of a source of the speech input signal.

19. The speech dialog system of claim 1 , where the analysis signal comprises a real number between approximately zero and approximately one representing a probability measure for one of the one or more non-semantic characteristics of the speech input signal.

20. A speech dialog system, comprising:

a processor;

a signal pre-processor unit configured to process a speech input signal and generate an enhanced speech signal;

a speech recognition unit configured to receive the enhanced speech signal and generate a recognition result containing one or more words spoken in the speech input signal based on the enhanced speech signal; and

a speech dialog control unit configured to receive the recognition result, the speech dialog control unit configured to control the signal pre-processor unit based upon the recognition result.

21. The speech dialog system of claim 20 , where the signal pre-processor unit comprises a noise reduction filter or an echo compensation filter; and

where the speech dialog control unit is configured to adjust one or more parameters of the noise reduction filter or the-echo compensation filter of the pre-processor unit based on the recognition result.

22. A method, comprising the steps of:

performing, via a processor, operations of:

processing a speech input signal to generate an enhanced speech signal;

analyzing the speech input signal or the enhanced speech signal to generate an analysis signal that comprises information related to one or more non-semantic characteristics of the speech input signal;

generating a recognition result containing one or more words spoken in the speech input signal based on the enhanced speech signal;

output, via a speech output unit, a synthesized speech output in response to the recognition result;

controlling the speech output unit based on the information of the analysis signal; and

adjusting one or more parameters used to process the speech input signal based on the recognition result.

23. A computer program product comprising a non-transitory computer readable medium having computer executable instructions executable by a computer for controls based on a speech dialog, the computer program product comprising:

computer code for processing a speech input signal to generate an enhanced speech signal;

computer code for analyzing the speech input signal or the enhanced speech signal to generate an analysis signal that comprises information related to one or more non-semantic characteristics of the speech input signal;

computer code for generating a recognition result containing one or more words spoken in the speech input signal based on the enhanced speech signal;

computer code for outputting, via a speech output unit, a synthesized speech output in response to the recognition result;

computer code for controlling the speech output unit based on the information of the analysis signal; and

computer code for adjusting one or more parameters used to process the speech input signal based on the recognition result.

Assignments (11)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSET PURCHASE AGREEMENT Recorded Jan 19, 2010
From: HARMAN BECKER AUTOMOTIVE SYSTEMS GMBH
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 023810/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2008
From: LOEW, ANDREAS
To: HARMAN BECKER AUTOMOTIVE SYSTEMS GMBH
Reel/Frame 020851/0131 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2008
From: SCHMIDT, GERHARD UWE
To: HARMAN BECKER AUTOMOTIVE SYSTEMS GMBH
Reel/Frame 020850/0972 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2008
From: KOENIG, LARS
To: HARMAN BECKER AUTOMOTIVE SYSTEMS GMBH
Reel/Frame 020850/0896 →
Priority Claims (1)
EP 06025974 · Dec 14, 2006 · regional
Continuity (1)
Related Publication 20080147397A1 · Jun 19, 2008