IP Library Granted Patent US 7,734,472
Granted Patent B2
US 7,734,472 · App. 10/951,594 · Granted Jun 8, 2010

Speech recognition enhancer

Assignee: Alcatel
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,734,472
App. No.
10/951,594
Granted
Jun 8, 2010
Kind
B2
Abstract

The invention concerns a speech recognition enhancer ( 51 ) and a speech recognition system comprising such speech recognition enhancer ( 51 ), an audio input unit ( 41 ) and a speech recognizer ( 61, 3 ). The speech recognition enhancer ( 51 ) is arranged between the audio input unit ( 41 ) and the speech recognizer ( 61, 3 ). The speech recognition enhancer ( 51 ) has a parametrizable pre-filtering unit ( 511 ), a parametrizable dynamic voice level control unit ( 512 ), a parametrizable noise reduction unit ( 513 ) and a parametrizable voice level control unit ( 514 ). The parameters of these parametrizable units ( 511, 512, 513, 514 ) are adjusted to the characteristics of the specific audio input unit ( 41 ) and/or the characteristics of the specific speech recognizer ( 61, 3 ) for adapting the audio input unit ( 41 ) to the speech recognizer ( 61, 3 ).

Claims (47)

1. A speech recognition system comprising:

an audio input unit arranged in a terminal;

a speech recognizer;

an adjustable speech recognition enhancer, separate from and arranged between the audio input unit and the speech recognizer, comprising:

a parametrizable pre-filtering unit,

a parametrizable dynamic voice level control unit,

a parametrizable noise reduction unit,

a parametrizable voice level control unit, and

a parameter setting unit;

wherein the parameters of the parametrizable pre-filtering unit, the parametrizable dynamic voice level control unit, the parametrizable noise reduction unit and the parametrizable voice level control unit are adjusted by the parameter setting unit to the characteristics of the specific audio input unit and/or the characteristics of the specific speech recognizer of the speech recognition system for adapting the audio input unit to the speech recognizer,

wherein parameters set by the parameter setting unit are stored in a data storage, and

wherein the adjustable speech recognition enhancer is adapted to process speech wave forms only in the time domain.

2. The speech recognition system of claim 1 ,

wherein the speech recognition system is a distributed speech recognition system,

wherein the speech recognizer comprises:

a central speech recognition server and at least one remote distributed speech recognition front-end performing the process of feature extraction,

wherein the distributed speech recognition front-end is located in a respective terminal,

wherein the terminal comprises a respective speech recognition enhancer and a respective audio input unit,

wherein each speech recognition enhancer is arranged in-between the audio input unit and the distributed speech recognition front-end of the respective terminal, and

wherein the parameters of the parametrizable pre-filtering unit, the parametrizable dynamic voice level control unit, the parametrizable noise reduction unit and the parametrizable voice level control unit of each speech recognition enhancer are adjusted to the characteristics of the respective audio input unit and/or the characteristics of the respective distributed speech recognition front-end of the respective terminal for adapting the audio input unit of this terminal to the distributed speech recognition front-end of this terminal.

3. The speech recognition system of claim 1 ,

wherein the speech recognition system is a stand-alone speech recognition system, wherein the speech recognizer is embedded in the terminal.

4. The speech recognition system of claim 1 ,

wherein the audio input unit comprises a microphone and an analogue to digital converter, and

wherein the speech recognizer is adapted to perform a transformation of a speech signal from the time domain to the frequency domain and to perform further processing of the transformed speech signal in the frequency domain.

5. The speech recognition system of claim 1 ,

wherein the pre-filtering unit is adapted to perform high-pass filtering of the speech signal provided by the audio input unit, and

wherein the parameters of the pre-filtering unit are adjusted to the characteristics of the audio input unit, in particular to the characteristics of a microphone arranged in the terminal.

6. The speech recognition system of claim 1 ,

wherein the dynamic voice level control unit is adapted to provide a dynamic voice level compression of the output signal of the pre-filtering unit, the dynamic voice level compression depending on parameters specifying a compression factor and a nominal voice level.

7. The speech recognition system of claim 1 ,

wherein the noise reduction unit comprises a voice activity detector and an amplifier controlled by the voice activity detector, and

wherein the voice activity detector is adapted to reduce the amplification factor of the amplifier when detecting a speech pause.

8. The speech recognition system of claim 1 ,

wherein the voice level control unit contains an amplifier for adapting the voice level of the output signal of the noise reduction unit to a preset voice level adapted to the characteristics of the speech recognizer.

9. A speech recognition enhancer for arrangement separate from and between an audio input unit and a speech recognizer of a speech recognition system comprising

a parametrizable pre-filtering unit,

a parametrizable dynamic voice level control unit,

a parametrizable dynamic voice level control unit,

a parametrizable noise reduction unit, and

a parametrizable voice level control unit,

wherein the parameters of the parametrizable pre-filtering unit, the parametrizable dynamic voice level control unit, the parametrizable noise reduction unit and the parametrizable voice level control unit are adjustable by a parameter setting unit to the characteristics of the specific audio input unit and/or the characteristics of the specific speech recognizer for adapting the audio input unit to the speech recognizer,

wherein parameters set by the parameter setting unit are stored in a data storage,

wherein the speech recognition enhancer is adapted to process speech wave forms only in the time domain.

10. The speech recognition system of claim 2 , wherein the central speech recognition server and the at least one remote distributed speech recognition front-end communicate via a communication network.

11. The speech recognition system of claim 10 , wherein the communication network is one of a Global System for Mobile Communications (GSM) network, a Universal Telecommunications System (UMTS) network or a DCMA 2000 network.

12. The speech recognition system of claim 10 , wherein the communication network is one of Public Switched Telecommunication Network (PSTN) or Integrated Service Digital Network (ISDN).

Assignments (9)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: PROVENANCE ASSET GROUP LLC
To: RPX CORPORATION
Reel/Frame 059352/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: CORTLAND CAPITAL MARKETS SERVICES LLC
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058983/0104 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: NOKIA US HOLDINGS INC.
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058363/0723 →
ASSIGNMENT AND ASSUMPTION AGREEMENT Recorded Feb 14, 2019
From: NOKIA USA INC.
To: NOKIA US HOLDINGS INC.
Reel/Frame 048370/0682 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP LLC
To: NOKIA USA INC.
Reel/Frame 043879/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP, LLC
To: CORTLAND CAPITAL MARKET SERVICES, LLC
Reel/Frame 043967/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2017
From: NOKIA TECHNOLOGIES OY; NOKIA SOLUTIONS AND NETWORKS BV; ALCATEL LUCENT SAS
To: PROVENANCE ASSET GROUP LLC
Reel/Frame 043877/0001 →
CHANGE OF NAME Recorded Aug 14, 2014
From: ALCATEL
To: ALCATEL LUCENT
Reel/Frame 033542/0605 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2004
From: WALKER, MICHAEL
To: ALCATEL
Reel/Frame 015848/0053 →
Priority Claims (1)
EP 03292957 · Nov 27, 2003 · regional
Continuity (1)
Related Publication 20050119886A1 · Jun 2, 2005