IP Library › Granted Patent US 7,676,363
Granted Patent B2
US 7,676,363 · App. 11/427,590 · Granted Mar 9, 2010

Automated speech recognition using normalized in-vehicle speech

Assignee: General Motors LLC
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,676,363
App. No.
11/427,590
Filed
Jun 29, 2006
Granted
Mar 9, 2010
Kind
B2
Art Unit
2626
USPC
704/234
Abstract

A speech recognition method includes the steps of receiving speech in a vehicle, extracting acoustic data from the received speech, and applying a vehicle-specific inverse impulse response function to the extracted acoustic data to produce normalized acoustic data. The speech recognition method may also include one or more of the following steps: pre-processing the normalized acoustic data to extract acoustic feature vectors; decoding the normalized acoustic feature vectors using as input at least one of a plurality of global acoustic models built according to a plurality of Lombard levels of a Lombard speech corpus covering a plurality of vehicles; calculating the Lombard level of vehicle noise; and/or selecting the at least one of the plurality of global acoustic models that corresponds to the calculated Lombard level for application during the decoding step.

Claims (20)

1. A method of speech recognition comprising the steps of: (a) receiving speech in a vehicle; (b) extracting acoustic data from the received speech; and (c) applying a vehicle-specific inverse impulse response function to the extracted acoustic data to produce normalized acoustic data.

2. The method of claim 1 , further comprising the steps of: (d) pre-processing the normalized acoustic data to extract normalized acoustic feature vectors; and (e) decoding the normalized acoustic feature vectors using as input at least one of a plurality of global acoustic models, wherein each model is distinguished from the other models based on a Lombard level.

3. The method of claim 2 , further comprising the steps of: (f) calculating the Lombard level of vehicle noise; and (g) selecting the global acoustic model of the plurality of global acoustic models that corresponds to the calculated Lombard level for application during the decoding step (e).

4. The method of claim 2 , wherein the global acoustic models are based on a Lombard speech corpus developed in a sound-controlled environment using a plurality of speakers, a plurality of different levels of noises, and a plurality of different utterances.

5. The method of claim 4 , wherein the global acoustic models are built after extracting acoustic features of speech from the Lombard speech corpus without noise reduction.

6. The method of claim 4 , wherein the global acoustic models are built by selecting a Lombard level, gathering speech data corresponding to the selected Lombard level from the Lombard speech corpus, selecting one or more vehicles that exhibit noise at the selected Lombard level, mixing the gathered speech data with acoustic data from the selected one or more vehicles, and extracting acoustic features of speech from the Lombard speech corpus with noise reduction.

7. The method of claim 2 , wherein the decoding step is performed in-vehicle.

8. The method of claim 2 , wherein the decoding step is performed in a remote server.

9. The method of claim 1 , wherein the inverse impulse response function is determined by first determining an impulse response function for the vehicle and then mathematically calculating the inverse of the impulse response function.

10. The method of claim 1 , wherein the inverse impulse response function is determined by determining correlation and/or covariance between an audio signal received from an integrated vehicle microphone (IVM) on a first channel and an audio signal received from a mouth reference position (MRP) microphone at a second channel, wherein the IVM is designated as an input and the MRP microphone is designated as an output.

11. A method of speech recognition for a plurality of vehicles, comprising the steps of: (a) developing a corpus of Lombard speech data; (b) building a plurality of global acoustic models based on the corpus of Lombard speech data; (c) receiving speech in a vehicle using an integrated vehicle microphone; (d) generating an inverse impulse response function for each of the plurality of vehicles; (e) extracting acoustic data from the received speech; and (f) applying the vehicle-specific inverse impulse response function to the extracted acoustic data to produce normalized acoustic data.

12. The method of claim 11 , wherein the corpus of Lombard speech data includes a plurality of Lombard levels, and wherein each model of the plurality of global acoustic models is distinguished from the other models based on a Lombard level of the plurality of Lombard levels.

13. The method of claim 11 , wherein the inverse impulse response function is determined by first determining an impulse response function for the vehicle and then mathematically calculating the inverse of the impulse response function.

14. The method of claim 11 , wherein the inverse impulse response function is determined by determining correlation and/or covariance between an audio signal received from an integrated vehicle microphone (IVM) on a first channel and an audio signal received from a mouth reference position (MRP) microphone at a second channel, wherein the IVM is designated as an input and the MRP microphone is designated as an output.

15. The method of claim 11 , wherein the Lombard speech corpus is developed in a sound-controlled environment using a plurality of speakers, a plurality of different levels of noises, and a plurality of different utterances.

16. The method of claim 15 , wherein the global acoustic models are built after extracting acoustic features of speech from the Lombard speech corpus without noise reduction.

17. The method of claim 16 , wherein the global acoustic models are built by selecting a Lombard level, gathering speech data corresponding to the selected Lombard level from the Lombard speech corpus, selecting one or more vehicles that exhibit noise at the selected Lombard level, mixing the gathered speech data with acoustic data from the selected one or more vehicles, and extracting acoustic features of speech from the Lombard speech corpus with noise reduction.

18. The method of claim 11 , further comprising the steps of: (g) pre-processing the normalized acoustic data to extract acoustic feature vectors; (h) decoding the normalized acoustic feature vectors using as input at least one of a plurality of global acoustic models, wherein each model is distinguished from the other models based on a Lombard level of a Lombard speech corpus covering a plurality of vehicles; (i) calculating the Lombard level of vehicle noise; and (j) selecting the at least one of the plurality of global acoustic models that corresponds to the calculated Lombard level for application during the decoding step (h).

19. The method of claim 18 wherein the decoding step is performed in a remote server.

20. A method of speech recognition for a plurality of vehicles, comprising the steps of: (a) developing a corpus of Lombard speech including a plurality of Lombard levels; (b) building a plurality of global acoustic models based on the corpus of Lombard speech data, wherein each model of the plurality of global acoustic models is distinguished from the other models based on a Lombard level of the plurality of Lombard levels; (c) receiving speech in a vehicle using an integrated vehicle microphone; (d) generating an inverse impulse response function for each of the plurality of vehicles, wherein the inverse impulse response function is determined by at least one of first determining an impulse response function for the vehicle and then mathematically calculating the inverse of the impulse response function, or determining correlation and/or covariance between an audio signal received from an integrated vehicle microphone (IVM) on a first channel and an audio signal received from a mouth reference position (MRP) microphone at a second channel wherein the IVM is designated as an input and the MRP microphone is designated as an output; (e) extracting acoustic data from the received speech; (f) applying the vehicle-specific inverse impulse response function to the extracted acoustic data to produce normalized acoustic data; (g) pre-processing the normalized acoustic data to extract acoustic feature vectors; (h) decoding the normalized acoustic feature vectors using as input at least one of a plurality of global acoustic models, wherein each model is distinguished from the other models based on a Lombard level of a Lombard speech corpus covering a plurality of vehicles; (i) calculating the Lombard level of vehicle noise; and (j) selecting the at least one of the plurality of global acoustic models that corresponds to the calculated Lombard level for application during the decoding step (h).

Assignments (14)
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2014
From: WILMINGTON TRUST COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 034183/0436 →
SECURITY AGREEMENT Recorded Nov 8, 2010
From: GENERAL MOTORS LLC
To: WILMINGTON TRUST COMPANY
Reel/Frame 025327/0196 →
RELEASE OF SECURITY INTEREST Recorded Nov 5, 2010
From: UAW RETIREE MEDICAL BENEFITS TRUST
To: GENERAL MOTORS LLC
Reel/Frame 025315/0162 →
RELEASE OF SECURITY INTEREST Recorded Nov 4, 2010
From: UNITED STATES DEPARTMENT OF THE TREASURY
To: GM GLOBAL TECHNOLOGY OPERATIONS, INC.
Reel/Frame 025245/0587 →
CHANGE OF NAME Recorded Nov 12, 2009
From: GENERAL MOTORS COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 023504/0691 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2009
From: MOTORS LIQUIDATION COMPANY
To: GENERAL MOTORS COMPANY
Reel/Frame 023148/0248 →
SECURITY AGREEMENT Recorded Aug 27, 2009
From: GENERAL MOTORS COMPANY
To: UNITED STATES DEPARTMENT OF THE TREASURY
Reel/Frame 023155/0814 →
SECURITY AGREEMENT Recorded Aug 27, 2009
From: GENERAL MOTORS COMPANY
To: UAW RETIREE MEDICAL BENEFITS TRUST
Reel/Frame 023155/0849 →
RELEASE OF SECURITY INTEREST Recorded Aug 21, 2009
From: CITICORP USA, INC. AS AGENT FOR BANK PRIORITY SECURED PARTIES; CITICORP USA, INC. AS AGENT FOR HEDGE PRIORITY SECURED PARTIES
To: MOTORS LIQUIDATION COMPANY (F/K/A GENERAL MOTORS CORPORATION)
Reel/Frame 023119/0817 →
CHANGE OF NAME Recorded Aug 21, 2009
From: GENERAL MOTORS CORPORATION
To: MOTORS LIQUIDATION COMPANY
Reel/Frame 023129/0236 →
RELEASE OF SECURITY INTEREST Recorded Aug 20, 2009
From: UNITED STATES DEPARTMENT OF THE TREASURY
To: MOTORS LIQUIDATION COMPANY (F/K/A GENERAL MOTORS CORPORATION)
Reel/Frame 023119/0491 →
SECURITY AGREEMENT Recorded Apr 16, 2009
From: GENERAL MOTORS CORPORATION
To: CITICORP USA, INC. AS AGENT FOR BANK PRIORITY SECURED PARTIES; CITICORP USA, INC. AS AGENT FOR HEDGE PRIORITY SECURED PARTIES
Reel/Frame 022552/0006 →
SECURITY AGREEMENT Recorded Feb 3, 2009
From: GENERAL MOTORS CORPORATION
To: UNITED STATES DEPARTMENT OF THE TREASURY
Reel/Frame 022191/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2007
From: CHENGALVARAYAN, RATHINAVELU; PENNOCK, SCOTT M.
To: GENERAL MOTORS CORPORATION
Reel/Frame 018915/0900 →
Continuity (1)
Related Publication 20080004875A1 · Jan 3, 2008